Model training method, device and moving object re-identification method
By acquiring and analyzing the features of moving objects from multiple perspectives in the weakly supervised motion object re-identification model, determining the relationships between them and performing model training, the problem of lack of generalization matching rules in cross-view learning is solved, and the performance reliability and generalization ability of the model are improved.
Patent Information
- Application Number
- CN202110343626.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-30
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-03-30
AI Technical Summary
In the cross-view learning process of weakly supervised motion object re-identification model, the model performance is not reliable due to the lack of generalized cross-view matching rules.
By obtaining a subset of samples of moving objects from multiple perspectives, each sample subset contains the characteristics of moving objects from all moving objects from one perspective. The target learning model is used to analyze and determine the relationship between the features of moving objects from different perspectives, and then the model is trained to obtain the re-identification model of moving objects.
By learning the consistency of feature distributions from different perspectives, the algorithm's severe dependence on pseudo-labels is alleviated, and the generalization and performance reliability of the model are improved.
Smart Images

Figure CN115147453B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of machine learning, and more specifically, to a model training method, device, and moving object re-identification method. Background Art
[0002] Moving object re-identification studies how to match moving objects under different cameras. This technology can be widely used in security, smart retail and other scenarios. According to the demand for data, moving object re-identification can be divided into three types: moving object re-identification based on fully supervised data, unsupervised moving object re-identification and weakly supervised moving object re-identification. Among them: moving object re-identification based on fully supervised data currently has better performance, but labeling a moving object picture requires comparison with images under all cameras, which is slow and costly, and difficult to collect large-scale data; although unsupervised moving object re-identification methods have made some progress, their performance is still poor and difficult to meet practical applications; moving object re-identification methods based on weakly supervised data, specifically, use data with only in-view labels to learn a moving object re-identification model. Weakly supervised data labeling is relatively simple, which is convenient for collecting large-scale data, and at the same time can bring certain supervisory signals to the model to ensure the performance of the model. It is currently the most commonly used method.
[0003] For the solution of learning the moving object re-identification model using weakly supervised data, the industry currently usually matches the moving object images under different cameras by feature similarity, and tries to make the model achieve better performance by bringing these samples closer. However, in practical applications, due to the large differences in image styles (lighting, background, etc.) between different cameras, matching directly based on feature similarity lacks reliable supervision signals, and the model is difficult to overcome the style changes between cameras, and the results are often unsatisfactory. Specifically, if the designed matching conditions are too strict, the successfully connected cross-view samples themselves have a high similarity, and learning how to bring them closer is of little value to the model; while loose matching conditions can easily produce incorrect matching results, which damages the performance of the model. This also leads to the fact that in the cross-view matching algorithm, moderate matching conditions are within a narrow parameter range, and the model generalization is poor.
[0004] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention
[0005] The embodiments of the present application provide a model training method, device and motion object re-identification method to at least solve the technical problem of low model performance reliability due to the lack of generalized cross-view matching rules in the cross-view learning process of the weakly supervised motion object re-identification model.
[0006] According to one aspect of an embodiment of the present application, a model training method is provided, including: obtaining sample subsets of moving objects under multiple perspectives, wherein each sample subset contains moving object features of all moving objects under one perspective, and the moving objects are objects that appear under the multiple perspectives during movement; inputting each of the sample subsets into a target learning model for analysis to obtain the moving object features under each perspective; determining the relationship between the moving object features under different perspectives based on the moving object features under each perspective; training the target learning model based on the relationship between the moving object features under each perspective and the moving object features under different perspectives to obtain a moving object re-identification model.
[0007] According to another aspect of an embodiment of the present application, a moving object re-identification method is also provided, including: obtaining a target moving object image to be identified; inputting the target moving object image into a moving object re-identification model for analysis to obtain an image recognition result; wherein the moving object re-identification model is obtained by obtaining sample subsets of moving objects under multiple perspectives, each sample subset containing moving object features of all moving objects under one perspective, the moving objects being objects that appear under the multiple perspectives during movement, each of the sample subsets is input into a target learning model for analysis to obtain moving object features under each perspective, determining the relationship between moving object features under different perspectives based on the moving object features under each perspective, and training the target learning model based on the relationship between the moving object features under each perspective and the moving object features under different perspectives.
[0008] According to another aspect of an embodiment of the present application, a model training device is also provided, including: an acquisition module, used to acquire sample subsets of moving objects under multiple perspectives, wherein each sample subset contains moving object features of all moving objects under one perspective; an analysis module, used to input each of the sample subsets into a target learning model for analysis, to obtain the moving object features under each perspective; a determination module, used to determine the relationship between the moving object features under different perspectives based on the moving object features under each perspective; a training module, used to train the target learning model based on the relationship between the moving object features under each perspective and the moving object features under different perspectives, to obtain a moving object re-identification model.
[0009] According to another aspect of an embodiment of the present application, a non-volatile storage medium is also provided, wherein the non-volatile storage medium includes a stored program, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the above-mentioned model training method.
[0010] According to another aspect of an embodiment of the present application, an electronic device is also provided, including: a processor and a memory connected to the processor, the memory being used to provide the processor with instructions for processing the following processing steps: obtaining sample subsets of moving objects under multiple perspectives, wherein each sample subset contains moving object features of all moving objects under one perspective, and the moving objects are objects that appear under the multiple perspectives during movement; inputting each of the sample subsets into a target learning model for analysis to obtain moving object features under each perspective; determining the relationship between moving object features under different perspectives based on the moving object features under each perspective; training the target learning model based on the relationship between the moving object features under each perspective and the moving object features under different perspectives to obtain a moving object re-identification model.
[0011] In an embodiment of the present application, the problem of weakly supervised moving object re-identification is solved by learning the consistency of feature distribution under different perspectives, thereby alleviating the previous algorithm's heavy reliance on pseudo-labels. Specifically, based on weakly supervised data, a feature extraction network is trained so that the network can recognize different moving object images under the same perspective, and learn a discriminative moving object feature for each moving object under each perspective; through the consistency constraint of sample relationships under different perspectives, the moving object feature knowledge learned based on the label information within the perspective is passed to other perspectives, and the moving object feature knowledge that is perspective-invariant and sufficiently generalizable is obtained, so that the model can recognize moving object images under different perspectives. The present application solves the technical problem that in the cross-perspective learning process of the weakly supervised moving object re-identification model, the model performance reliability is not high due to the lack of generalized cross-perspective matching rules. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0013] Figure 1 is a schematic diagram of the hardware structure of a computer terminal according to Embodiment 1 of the present application;
[0014] Figure 2 is a flow chart of a model training method according to Example 1 of the present application;
[0015] Figure 3 is a flow chart of a moving object re-identification method according to Embodiment 2 of the present application;
[0016] Figure 4a is a schematic diagram of an optional pedestrian re-identification process according to Embodiment 2 of the present application;
[0017] Figure 4b is a schematic diagram of another optional pedestrian re-identification process according to Embodiment 2 of the present application;
[0018] Figure 5 It is a structural schematic diagram of a model training device according to Example 3 of the present application. DETAILED DESCRIPTION
[0019] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.
[0020] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0021] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following explanations:
[0022] Moving object re-identification: also known as moving object re-recognition, is a technology that uses computer vision technology to determine whether there is a specific moving object in an image or video sequence. It is widely considered to be a sub-problem of image retrieval. Moving object re-identification aims to learn a model that can extract identity features for moving object images, that is, to extract features that can fully express the appearance of the moving object. Based on this model, pictures belonging to the same moving object under different cameras can be matched together. By giving a moving object image, the moving object image under cross-device can be retrieved, which can make up for the visual limitations of fixed cameras and can be combined with moving object detection / moving object tracking technology. It can be widely used in smart security and other fields.
[0023] Weakly supervised model: A model learned using weakly supervised data.
[0024] Weakly labeled data / in-view labels: Weak supervision in this application refers to assigning the same label to multiple images of the same moving object under each camera, that is, the identity label of each image only represents the identity of the image under the current camera; in contrast, fully supervised data assigns the same label to multiple images of the same moving object under multiple cameras (camera network).
[0025] Cross-view / cross-camera samples: samples taken from different cameras.
[0026] Same view / same camera samples: samples taken from the same camera.
[0027] Cross-view learning: enables the moving object re-identification model to recognize moving objects under different cameras.
[0028] Softmax function: also known as the normalized exponential function, it is a generalization of the binary classification function sigmoid in multi-classification. It can compress a K-dimensional vector z containing any real number into another K-dimensional real vector σ(z), so that the range of each element is between (0, 1), and the sum of all elements is 1. This function has a very wide range of applications in machine learning and deep learning, especially in dealing with multi-classification (C>2) problems. The final output unit of the classifier needs to be numerically processed by the Softmax function. The definition of the Softmax function is: Among them, Vi is the output of the previous output unit of the classifier, i represents the category index, C is the total number of categories, S i It represents the ratio of the exponent of the current element to the sum of the exponents of all elements. Softmax converts the output values of multiple categories into relative probabilities, making it easier to understand and compare.
[0029] Kullback-Leibler divergence (KL divergence): also known as relative entropy, if there are two separate probability distributions P(x) and Q(x) for the same random variable X, KL divergence can be used to measure the difference between the two probability distributions. In machine learning, P(x) is often used to represent the true distribution of samples, and Q(x) is used to represent the distribution predicted by the model. The smaller the KL divergence, the closer the distributions of P(x) and Q(x) are. The distribution of Q(x) can be made close to P(x) by repeatedly training Q(x).
[0030] Cross entropy: Break down the KL divergence formula:
[0031]
[0032] The former H(p(x)) represents information entropy, and the latter is cross entropy, that is, KL divergence = cross entropy - information entropy. The cross entropy formula is expressed as: When training a network for machine learning, since the input data and labels are often already determined, the true probability distribution P(x) is also determined, so the information entropy is a constant here. Since the value of KL divergence represents the difference between the true probability distribution P(x) and the predicted probability distribution Q(x), the smaller the value, the better the prediction result, so it is necessary to minimize the KL divergence, and the cross entropy is equal to the KL divergence plus a constant (information entropy), and the formula is easier to calculate than the KL divergence, so the cross entropy loss function is often used to calculate the loss function in machine learning.
[0033] Euclidean distance: Euclidean distance is a commonly used distance definition, which refers to the true distance between two points in m-dimensional space, or the natural length of a vector (that is, the distance from the point to the origin).
[0034] Jensen-Shannon (JS divergence): also known as JS distance, is a variation of KL divergence. The difference between it and JS divergence is that the range of JS divergence is [0,1], which is 0 for the same and 1 for the opposite. Compared with KL, it is more accurate in judging similarity. At the same time, JS divergence is symmetric, JS(P||Q)=JS(Q||P), and its formula is:
[0035]
[0036] Mean square error loss function: commonly used in the least squares method, its idea is to minimize the distance from each training point to the fitting line (minimize the sum of squares). Its definition expressed in the form of an activation function in a neural network is as follows:
[0037]
[0038] Among them, a=f(z)=f(w·x+b): x is the input, w and b are the parameters of the network, and f(*) is the activation function.
[0039] Smooth L1 loss function: In deep neural networks or recurrent neural networks, the gradient of the error will be accumulated and multiplied during the update. If the gradient value between network layers is greater than 1, repeated multiplication will cause the gradient to grow exponentially and become very large, which will then cause the network weights to be updated significantly, making the network unstable, i.e., gradient explosion. To solve this problem, the Smooth L1 loss function is proposed. For regression problems in target detection, since the mean square error loss ||yf(z)|| is often used, 2, so that when back propagation is used to derive w or b, yf(z) still exists. Then, when the predicted value and the target value differ greatly, it is easy to cause gradient explosion. Therefore, ||yf(z)|| 2 This form of mean square error is transformed into smooth L1 (yf(z)) in the form of
[0040]
[0041] When |yf(z)|<1, it is not easy to cause gradient explosion. At this time, it is restored to the mean square error loss form and a smoothing coefficient of 0.5 is given, that is, 0.5||yf(z)|| 2 ; When |yf(z)|≥1, it is easy to cause gradient explosion. At this time, the loss power is reduced to |yf(z)|-0.5. At this time, the term yf(z) does not exist during the back propagation derivation, thus preventing the gradient explosion.
[0042] Example 1
[0043] According to an embodiment of the present application, an embodiment of a model training method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0044] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 FIG. 1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing a model training method. Figure 1 As shown, the computer terminal 10 (or mobile device 10) may include one or more (102a, 102b, ..., 102n are used to illustrate) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. For example, the computer terminal 10 may also include Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.
[0045] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuits". The data processing circuits may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuit acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0046] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the model training method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, realizing the vulnerability detection method of the above-mentioned application. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0047] The transmission module 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission module 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet; alternatively, the transmission module 106 can be a radio frequency (Radio Frequency, RF) module, which is used to communicate with the Internet wirelessly.
[0048] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).
[0049] In the above operating environment, the embodiment of the present application provides a model training method for moving object re-identification, such as Figure 2 As shown, the specific process of the method includes steps S202-S208, wherein:
[0050] Step S202 , obtaining sample subsets of moving objects under multiple viewing angles, wherein each sample subset contains moving object features of all moving objects under one viewing angle, and the moving objects are objects that appear in sequence under multiple viewing angles during movement.
[0051] In an optional embodiment of the present application, samples of multiple moving objects captured by multiple cameras (i.e., from multiple perspectives) are obtained as a moving object sample data set, and the samples may be images or videos of the moving objects. Typically, each moving object sample corresponds to a moving object feature. Then, the moving object samples captured by different cameras are divided, and all moving object samples captured by the same camera are taken as a sample subset for training the target learning model.
[0052] It should be noted that the above-mentioned moving objects are not limited to pedestrians, but may also be other objects that appear in multiple shots in sequence, such as pets on the street, animals in a zoo, etc.
[0053] Step S204: input each sample subset into the target learning model for analysis to obtain the moving object features at each viewing angle.
[0054] In an optional embodiment of the present application, for any sample subset, all the motion object samples in the sample subset are input into the feature extraction network in the target learning model for analysis to obtain the motion object features corresponding to each motion object sample, wherein the feature extraction network can select resnet (Residual Network), typically resnet50, resnet101, etc., to output the motion object features corresponding to each motion object sample.
[0055] Then, according to the motion object features corresponding to each motion object sample in the sample subset, the loss function under the perspective corresponding to the sample subset is determined, and the cross entropy loss function can be used here; the cross entropy loss functions under multiple perspectives are accumulated and summed to obtain the first loss function; finally, the target learning model is trained based on the first loss function.
[0056] Specifically, suppose that a total of n moving object images under different viewing angles are obtained as sample data sets, with x i,j represents the jth moving object image under the i-th viewing angle, y i,j For x i,j ID tag representing x i,j The yth from the ith perspective i,j moving objects, f(*) represents the feature extraction network, then the image x i,j The corresponding moving object feature is f(x i,j ), in g i,kRepresenting the features of the kth moving object under the ith perspective, the cross entropy loss function under the ith perspective can be obtained:
[0057]
[0058] Among them, exp(*) represents an exponential function with the natural constant e as the base, and τ1 is the temperature parameter. Its size has no direct relationship with the final accuracy of the model. The role of τ1 can be compared to the learning rate. If τ1 is set relatively large during training, the predicted probability distribution will be smoother and the loss function will be large. As the training progresses, τ1 can be made smaller. This process can be called cooling, which is similar to the simulated annealing algorithm. This is why τ1 is called the temperature parameter.
[0059] After determining the loss function under a single perspective, the in-perspective learning process can be performed under n perspectives at the same time to determine the total loss function (i.e., the first loss function) as:
[0060]
[0061] The target learning model is iteratively trained based on the first loss function until the model converges and fully learns the characteristics of the moving object at each perspective.
[0062] Step S206: determining the relationship between the features of the moving object at different viewing angles according to the features of the moving object at each viewing angle.
[0063] In an optional embodiment of the present application, first, based on the moving object features at each viewing angle, a second loss function is determined when transferring the moving object features at multiple viewing angles; then, based on the moving object features at each viewing angle, a third loss function is determined when aligning the moving object feature distributions at multiple viewing angles; finally, based on the first loss function, the second loss function and the third loss function, the relationship between the moving object features at different viewing angles is determined.
[0064] In an optional embodiment of the present application, when determining the second loss function, the calculation may be performed according to the following steps:
[0065] Step S2061, mapping the first sample under the first perspective to the second perspective to obtain a pseudo label of the first sample under the second perspective, wherein the first perspective and the second perspective are any two different perspectives under multiple perspectives, the first sample is any moving object sample in the sample subset corresponding to the first perspective, and the pseudo label is the first similarity vector mapped to the first sample under the second perspective.
[0066] Specifically, the first sample is compared with the moving object feature under the second perspective to obtain a third similarity vector mapped from the first sample to the second perspective; for example, when mapping the sample under the i-th perspective to the z-th perspective, by i,j and the moving object features at the z-th perspective Compare (C z represents the number of moving objects under the z-th viewing angle), and the sample x i,j Mapped to the z-th perspective, we get the third similarity vector In this process, (*) represents the mapping function, then (x i,j , G z )The qth element is defined as follows:
[0067]
[0068] In order to make the result more accurate, the third similarity vector Adjustment is made, such as substituting the moving object features under the second perspective into the softmax function for calculation, obtaining the first relationship matrix between the target objects under the second perspective, and adjusting the third similarity vector according to the first relationship matrix to obtain the fourth similarity vector; for example, defining the relationship matrix between the moving objects in the z-th space Among them, π(.τ2) is the softmax function with temperature parameter τ2, which is applied in For each row of , the adjusted fourth similarity vector is:
[0069]
[0070] It is understandable that in order to improve the performance of the model, when transferring the moving object features under two perspectives, it is necessary to pay attention to and select those moving objects that may appear at the same time under the two perspectives. Therefore, by determining the effective moving object features under the second perspective, the fourth similarity vector can be further adjusted according to the effective moving object features, and finally the pseudo-label of the first sample under the second perspective is obtained. Specifically, when determining the effective moving object features under the second perspective, for any moving object feature under the second perspective, the moving object feature is compared with all the moving object features under the first perspective to obtain a likelihood vector; the likelihood vector is compared with a preset threshold, and if the likelihood vector is greater than the preset threshold, the moving object feature corresponding to the likelihood vector is determined to be a valid moving object feature. For example, each moving object feature g under the zth perspective z,k (1≤k≤C z ) and the moving object feature g under the i-th perspective i,k Compare and get the kth moving object feature g under the zth perspective z,k The likelihood vector in The value of the qth element of is:
[0071]
[0072] If and only if When the maximum value of is greater than the preset threshold θ, it is considered that the moving object g z,k If it appears at the i-th viewing angle, then g needs to be considered when mapping the sample at the i-th viewing angle to the z-th viewing angle. z,k . C at the zth viewing angle z After calculating the likelihood vector for each moving object, use represents the selected motion object features (i.e., effective motion object features), where If and only if Finally, the sample x i,j Pseudo label at the z-th view Will be based on get:
[0073]
[0074] Step S2062: self-map the first sample to the first viewing angle to obtain a soft label of the first sample at the first viewing angle, wherein the soft label is a second similarity vector of the first sample self-mapped to the first viewing angle.
[0075] Although according to the sample x i,j The feature information of the moving object at the i-th viewing angle can be directly obtained, but the information is embodied in the form of one-hot encoding, and the amount of information contained in the relationship between samples obtained from it is too small. Therefore, in an optional embodiment of the present application, the sample x i,j Mapping from itself to the i-th perspective, we get the sample x i,j Soft label at the i-th view The calculation process is the same as that in step S2061 The process is similar to
[0076]
[0077] in, It represents the moving object features (i.e., effective moving object features) selected when mapping the samples under the z-th viewing angle to the i-th viewing angle.
[0078] Step S2063: determining a relationship between sample groups according to the pseudo labels and the soft labels, wherein the sample group includes at least two samples under the first viewing angle, and the at least two samples include the first sample.
[0079] In an optional embodiment of the present application, at least two samples are selected from the first perspective to form a sample group, usually two or three samples are selected, and the similarity between these samples is calculated. The sample group should include the first sample.
[0080] Specifically, when a sample group includes two samples, a bigram relationship between the sample groups is determined based on the pseudo-label and the soft-label, and the bigram relationship is used to indicate the similarity between the two samples, wherein the bigram relationship includes one of the following relationships: KL divergence, Euclidean distance, and JS divergence.
[0081] In an optional embodiment of the present application, a second sample is selected from the first perspective to form a sample group with the first sample, wherein the second sample is any sample other than the first sample from the first perspective; at the first perspective, the KL divergence between the first sample and the second sample is calculated based on the soft label; at the second perspective, the KL divergence between the first sample and the second sample is calculated based on the pseudo label; for example, two samples x are selected from the i-th perspective i,j and x i,l , calculate the KL divergence of the two samples at the zth perspective respectively And the KL divergence of the two samples at the i-th perspective in:
[0082]
[0083]
[0084] Optionally, in addition to the above-mentioned KL divergence, other binary relationships between samples can also be defined. Specifically, a second sample and a first sample are selected from the first perspective to form a sample group, wherein the second sample is any sample other than the first sample from the first perspective, and a functional relationship such as the Euclidean distance or JS divergence between the sample groups is determined based on the pseudo-label and the soft label, wherein the Euclidean distance calculation formula between the sample groups is:
[0085]
[0086] The JS divergence calculation formula between sample groups is:
[0087]
[0088] When the sample group contains three samples, the triple relationship between the sample groups is determined based on the pseudo label and the soft label, and the triple relationship is used to indicate the similarity between the three samples. Specifically, the second sample and the third sample are selected from the first perspective to form a sample group with the first sample, wherein the second sample and the third sample are any two different samples except the first sample from the first perspective, and the triple relationship between the sample groups is determined based on the pseudo label and the soft label. For example, three samples x are selected from the i-th perspective. i,j , x i,l and x i,k , then the triple relationship between sample groups is:
[0089]
[0090] in,
[0091] Step S2064: determining a second loss function for transferring moving object features under multiple viewing angles according to the relationship between the sample groups.
[0092] In an optional embodiment of the present application, after determining the relationship between the sample groups, the obtained relationship between the sample groups can be substituted into the smooth L1 loss function or the mean square error loss function for calculation to obtain the loss function when the moving object feature is transferred between the first perspective and the second perspective; ψ(*) represents the smooth L1 loss function or the mean square error loss function, and the loss function when the moving object feature is transferred from the i-th perspective to the z-th perspective is:
[0093]
[0094] Among them, the mean square error loss function is:
[0095]
[0096] On this basis, the loss function of the moving object feature transfer between any two perspectives under all perspectives is calculated, and all the obtained loss functions are accumulated and summed to obtain the second loss function:
[0097]
[0098] In order to improve the performance of the model, the feature relationship of the same moving object under different viewing angles is further shortened by aligning the mean and variance of the moving object features under different viewing angles. In an optional embodiment of the present application, the third loss function is determined in the following manner: for any viewing angle, the first mean vector and the first standard deviation vector of all moving object features under the viewing angle are calculated, for example, μ is defined i , σ i is the feature of the moving object under the i-th perspective Calculate the second mean vector and the second standard deviation vector of all moving object features under all viewing angles, and define μ, σ as the moving object features under all viewing angles. The mean vector and standard deviation vector of i , the first standard deviation vector σ i , the second mean vector μ and the second standard deviation vector σ, determine the third loss function when aligning the feature distributions of moving objects under multiple perspectives:
[0099]
[0100] After obtaining the first loss function, the second loss function and the third loss function through the above process, the first loss function, the second loss function and the third loss function can be weighted summed to obtain the fourth loss function:
[0101] L=L SL +λ1*L CRL +λ2*L DA
[0102] Among them, λ1 and λ2 are weight values preset for the second loss function and the third loss function.
[0103] Based on the fourth loss function, the relationship between the features of the moving object under different perspectives can be determined for training the target learning model.
[0104] Step S208: training the target learning model according to the relationship between the moving object features at each viewing angle and the moving object features at different viewing angles to obtain a moving object re-identification model.
[0105] In an optional embodiment of the present application, the target learning model can be trained in the following manner:
[0106] Each time the data set is traversed, 8 images are randomly selected from each perspective for each training batch and substituted into the first loss function for calculation. The target learning model is trained by traversing the data set multiple times until convergence;
[0107] Each time the dataset is traversed, each training batch randomly selects 4 moving objects from each perspective, with 2 images for each moving object, that is, a total of 8 images from each perspective, and substituted them into the fourth loss function for calculation. The dataset is repeatedly traversed multiple times to train the target learning model until convergence, and finally a reliable moving object re-identification model is obtained.
[0108] In an embodiment of the present application, sample subsets of moving objects under multiple perspectives are obtained, wherein each sample subset contains moving object features of all moving objects under one perspective; each sample subset is input into a target learning model for analysis to obtain moving object features under each perspective; the relationship between moving object features under different perspectives is determined based on the moving object features under each perspective; the target learning model is trained based on the relationship between the moving object features under each perspective and the moving object features under different perspectives to obtain a moving object re-identification model. The present application solves the problem of weakly supervised moving object re-identification by learning the consistency of feature distribution under different perspectives, thereby solving the technical problem of low reliability of model performance due to the lack of generalized cross-perspective matching rules in the cross-perspective learning process of the weakly supervised moving object re-identification model.
[0109] Example 2
[0110] According to an embodiment of the present application, an embodiment of a method for re-identifying a moving object is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0111] Figure 3 is a flow chart of a moving object re-identification method provided in an embodiment of the present application, such as Figure 3 As shown, the method includes steps S302-S304, wherein:
[0112] Step S302: Acquire a target moving object image to be identified.
[0113] Step S304: input the target moving object image into a moving object re-recognition model for recognition to obtain an image recognition result.
[0114] Among them, the moving object re-identification model is obtained by obtaining sample subsets of moving objects under multiple perspectives, wherein each sample subset contains the moving object features of all moving objects under one perspective, and the moving object is an object that appears in multiple perspectives during the movement process. Each sample subset is input into the target learning model for training, and the moving object features under each perspective are determined. Based on the moving object features under each perspective, the relationship between the moving object features under different perspectives is determined, and the target learning model is trained based on the relationship between the moving object features under different perspectives.
[0115] Specifically, during the model training process, we first obtain samples of multiple moving objects captured by multiple cameras (i.e., from multiple perspectives) as a moving object sample data set. The samples can be images or videos of moving objects. Usually, each moving object sample corresponds to a moving object feature. Then, the moving object samples captured by different cameras are divided, and all moving object samples captured by the same camera are taken as a sample subset for training the target learning model.
[0116] For any sample subset, all the moving object samples in the sample subset are input into the feature extraction network in the target learning model for analysis to obtain the moving object features corresponding to each moving object sample, wherein the feature extraction network can select resnet50, resnet101, etc., to output the moving object features corresponding to each moving object sample; then, according to the moving object features corresponding to each moving object sample in the sample subset, the loss function under the perspective corresponding to the sample subset is determined, and the cross entropy loss function can be selected here; the cross entropy loss functions under multiple perspectives are accumulated and summed to obtain the first loss function; finally, the target learning model is trained based on the first loss function.
[0117] After determining the moving object features at each viewing angle, first determine the second loss function when transferring the moving object features at multiple viewing angles based on the moving object features at each viewing angle. Specifically, map the first sample at the first viewing angle to the second viewing angle to obtain the pseudo label of the first sample at the second viewing angle; self-map the first sample to the first viewing angle to obtain the soft label of the first sample at the first viewing angle; determine the relationship between the sample groups based on the pseudo label and the soft label; and determine the second loss function when transferring the moving object features at multiple viewing angles based on the relationship between the sample groups.
[0118] Then, according to the characteristics of the moving object at each perspective, the third loss function is determined when aligning the distribution of the moving object features at multiple perspectives; the first loss function, the second loss function and the third loss function are weighted and summed to obtain the fourth loss function, and the relationship between the moving object features at different perspectives is determined based on the fourth loss function; finally, the target learning model is trained according to the relationship between the moving object features at each perspective and the moving object features at different perspectives to obtain a moving object re-identification model.
[0119] It should be noted that since motion object re-identification has two application modes, motion object authentication and motion object retrieval, when applying the motion object re-identification model, it is necessary to determine the application mode of the current motion object re-identification model; wherein, the motion object authentication mode is used to determine whether the motion objects in multiple target motion object images are the same motion object, and the motion object retrieval mode is used to retrieve all images of the motion object in the target motion object image from a database, which includes images of multiple motion objects from multiple perspectives.
[0120] When it is determined that the current application mode is the moving object authentication mode, the moving object re-identification model analyzes the moving object features corresponding to the multiple target moving object images, determines the similarity between the multiple moving object features and compares them with a preset threshold, and determines whether the moving objects in the multiple target moving object images are the same moving object based on the comparison result; if the similarity between the multiple moving object features is greater than the preset threshold, it can be confirmed that the moving objects in the multiple target moving object images are the same moving object. Taking pedestrian re-identification as an example, an optional pedestrian re-identification process is as follows Figure 4a As shown, camera 1 captures pedestrian image 1, and camera 2 captures pedestrian image 2. The two cameras respectively send the captured pedestrian images to the pedestrian re-identification model running on the server. The pedestrian re-identification model analyzes and compares the pedestrian features corresponding to the two pedestrian images, and finally determines whether the pedestrians captured by camera 1 and camera 2 are the same pedestrian.
[0121] When it is determined that the current application mode is the motion object retrieval mode, the motion object re-identification model analyzes the motion object features corresponding to the target motion object image, and determines the similarity between the motion object features and the motion object features corresponding to all motion object images in the database, and compares them with a preset threshold, and determines all images of the motion object corresponding to the target motion object image in the database based on the comparison result. Specifically, all motion object images in the database whose feature similarity with the target motion object image is greater than the preset threshold can be considered as images of the motion object corresponding to the target motion object image. Taking pedestrian re-identification as an example, an optional pedestrian re-identification process is as follows: Figure 4b As shown, after camera 1 captures a pedestrian image, the pedestrian image is sent to a pedestrian re-identification model running on a server. The pedestrian re-identification model analyzes the pedestrian features corresponding to the pedestrian image and retrieves a database (which stores multiple pedestrian images captured by multiple cameras), performs feature comparison on the pedestrian image with all pedestrian images in the database, and finally outputs the image of the pedestrian under all cameras.
[0122] It should be noted that the moving objects in the embodiments of the present application are not limited to pedestrians, but may also be other objects that appear in multiple shots, such as pets on the street, animals in a zoo, etc. For example, in wildlife park management, given an animal image, the above model is used to find the images of the animal in all shots, so as to understand the activity range of the animal.
[0123] Example 3
[0124] According to an embodiment of the present application, a model training device for implementing the above-mentioned model training method is also provided, such as Figure 5 As shown, the device includes an acquisition module 50, an analysis module 52, a determination module 54 and a training module 56, wherein:
[0125] The acquisition module 50 is used to acquire sample subsets of moving objects under multiple viewing angles, wherein each sample subset contains moving object features of all moving objects under one viewing angle, and the moving objects are objects that appear under multiple viewing angles during movement.
[0126] In an optional embodiment of the present application, samples of multiple moving objects captured by multiple cameras are obtained as a moving object sample data set, and the samples can be images or videos of the moving objects. Typically, each moving object sample corresponds to a moving object feature; then, the moving object samples captured by different cameras are divided, and all moving object samples captured by the same camera are taken as a sample subset for training the target learning model.
[0127] The analysis module 52 is used to input each sample subset into the target learning model for analysis to obtain the characteristics of the moving object at each viewing angle.
[0128] In an optional embodiment of the present application, for any sample subset, all motion object samples in the sample subset are input into the feature extraction network in the target learning model for analysis to obtain the motion object features corresponding to each motion object sample, wherein the feature extraction network can select resnet50, resnet101, etc., to output the motion object features corresponding to each motion object sample; according to the motion object features corresponding to each motion object sample in the sample subset, a first loss function is determined; and the target learning model is trained based on the first loss function.
[0129] The determination module 54 is used to determine the relationship between the moving object features at different viewing angles according to the moving object features at each viewing angle.
[0130] In an optional embodiment of the present application, first, based on the moving object features at each viewing angle, a second loss function is determined when the moving object features are transferred under multiple viewing angles; specifically, the first sample under the first viewing angle is mapped to the second viewing angle to obtain a pseudo-label of the first sample under the second viewing angle; the first sample is self-mapped to the first viewing angle to obtain a soft label of the first sample under the first viewing angle; the relationship between the sample groups is determined based on the pseudo-labels and the soft labels; and the second loss function when the moving object features are transferred under multiple viewing angles is determined based on the relationship between the sample groups.
[0131] Then, based on the features of the moving object at each perspective, a third loss function is determined when aligning the distribution of the moving object features at multiple perspectives; the first loss function, the second loss function and the third loss function are weighted and summed to obtain a fourth loss function, and the relationship between the moving object features at different perspectives is determined based on the fourth loss function.
[0132] The training module 56 is used to train the target learning model according to the relationship between the moving object features at each viewing angle and the moving object features at different viewing angles to obtain a moving object re-identification model.
[0133] It should be noted that each module in the model training device in the embodiment of the present application corresponds one by one to the implementation steps of the model training method in Example 1. Since a detailed description has been given in Example 1, some details not reflected in this embodiment can be referred to Example 1 and will not be repeated here.
[0134] Example 4
[0135] According to an embodiment of the present application, an electronic device is also provided, which includes a processor and a memory, wherein: the memory is connected to the processor, and is used to provide the processor with instructions for processing the following processing steps: obtaining sample subsets of moving objects under multiple perspectives, wherein each sample subset contains moving object features of all moving objects under one perspective, and the moving object is an object that appears under multiple perspectives during movement; inputting each sample subset into a target learning model for analysis to obtain the moving object features under each perspective; determining the relationship between the moving object features under different perspectives based on the moving object features under each perspective; training the target learning model based on the relationship between the moving object features under each perspective and the moving object features under different perspectives to obtain a moving object re-identification model.
[0136] Optionally, the memory also stores instructions for processing the following steps: obtaining a target moving object image to be identified; inputting the target moving object image into a moving object re-identification model for analysis to obtain an image recognition result; wherein the moving object re-identification model is obtained by obtaining sample subsets of moving objects at multiple perspectives, each sample subset containing moving object features of all moving objects at one perspective, the moving object being an object that appears at multiple perspectives during movement, inputting each sample subset into a target learning model for analysis to obtain moving object features at each perspective, determining the relationship between moving object features at different perspectives based on the moving object features at each perspective, and training the target learning model based on the relationship between the moving object features at each perspective and the moving object features at different perspectives.
[0137] Since motion object re-identification has two application modes: motion object authentication and motion object retrieval, when applying the motion object re-identification model, it is necessary to determine the application mode of the current motion object re-identification model; wherein, the motion object authentication mode is used to determine whether the motion objects in multiple target motion object images are the same motion object, and the motion object retrieval mode is used to retrieve all images of the motion object in the target motion object image from a database, which includes images of multiple motion objects under multiple perspectives.
[0138] Example 5
[0139] According to an embodiment of the present application, a non-volatile storage medium is also provided, which includes a stored program, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the above-mentioned model training method.
[0140] Optionally, when the program is running, the device where the non-volatile storage medium is located is controlled to execute the following steps: obtaining sample subsets of moving objects under multiple perspectives, wherein each sample subset contains moving object features of all moving objects under one perspective, and the moving object is an object that appears under multiple perspectives during movement; inputting each sample subset into the target learning model for analysis to obtain the moving object features under each perspective; determining the relationship between the moving object features under different perspectives based on the moving object features under each perspective; training the target learning model based on the relationship between the moving object features under each perspective and the moving object features under different perspectives to obtain a moving object re-identification model.
[0141] Optionally, when the program is running, the device where the non-volatile storage medium is located is controlled to execute the following steps: obtaining a target moving object image to be identified; inputting the target moving object image into a moving object re-identification model for analysis to obtain an image recognition result; wherein the moving object re-identification model is obtained by obtaining sample subsets of moving objects under multiple perspectives, each sample subset containing moving object features of all moving objects under one perspective, and the moving object is an object that appears under multiple perspectives during movement, and each sample subset is input into a target learning model for analysis to obtain moving object features under each perspective, and based on the moving object features under each perspective, the relationship between the moving object features under different perspectives is determined, and the target learning model is trained based on the relationship between the moving object features under each perspective and the moving object features under different perspectives.
[0142] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0143] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0144] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0145] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0146] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0147] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or all or part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of each embodiment method of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk, etc. Various media that can store program codes.
[0148] The above are only preferred implementations of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A model training method, characterized in that: include: Acquire sample subsets of moving objects under multiple viewing angles, wherein each sample subset contains moving object features of all moving objects under one viewing angle, and the moving objects are objects that appear under the multiple viewing angles during movement; Input each of the sample subsets into the target learning model for analysis to obtain the features of the moving object at each viewing angle; Determining the relationship between the features of the moving object at different viewing angles according to the features of the moving object at each viewing angle; The target learning model is trained according to the relationship between the moving object features at each viewing angle and the moving object features at different viewing angles to obtain a moving object re-identification model; Among them, according to the motion object features at each viewing angle, the relationship between the motion object features at different viewing angles is determined, including: according to the motion object features corresponding to each motion object sample in each sample subset, the loss function at the viewing angle corresponding to each sample subset is determined, and the loss functions at multiple viewing angles are accumulated and summed to obtain a first loss function; according to the motion object features at each viewing angle, a second loss function is determined when the motion object features are transmitted at the multiple viewing angles; for any viewing angle, a first mean vector and a first standard deviation vector of all motion object features at the viewing angle are calculated, and a second mean vector and a second standard deviation vector of all motion object features at all viewing angles are calculated, and according to the first mean vector, the first standard deviation vector, the second mean vector and the second standard deviation vector, a third loss function is determined when the motion object feature distributions at the multiple viewing angles are aligned; according to the first loss function, the second loss function and the third loss function, the relationship between the motion object features at different viewing angles is determined.
2. The method according to claim 1, characterized in that Each of the sample subsets is input into the target learning model for analysis to obtain the features of the moving object at each viewing angle, including: For any sample subset, input the sample subset into a feature extraction network in a target learning model for analysis to obtain a motion object feature corresponding to each motion object sample in the sample subset; Determine a loss function at a viewing angle corresponding to the sample subset according to a moving object feature corresponding to each moving object sample in the sample subset; The loss functions under the multiple perspectives are accumulated and summed to obtain a first loss function; The target learning model is trained based on the first loss function.
3. The method according to claim 1, characterized in that Determining, according to the moving object feature at each viewing angle, a second loss function when transferring the moving object feature at the multiple viewing angles, comprising: Mapping a first sample under a first perspective to a second perspective to obtain a pseudo label of the first sample under the second perspective, wherein the first perspective and the second perspective are any two different perspectives among the multiple perspectives, the first sample is any moving object sample in a sample subset corresponding to the first perspective, and the pseudo label is a first similarity vector mapped from the first sample to the second perspective; Self-map the first sample to the first viewing angle to obtain a soft label of the first sample at the first viewing angle, wherein the soft label is a second similarity vector of the first sample self-mapped to the first viewing angle; Determining a relationship between a sample group according to the pseudo label and the soft label, wherein the sample group includes at least two samples under the first viewing angle, and the at least two samples include the first sample; A second loss function when transferring the moving object features under the multiple viewing angles is determined according to the relationship between the sample groups.
4. The method according to claim 3, characterized in that Mapping a first sample at a first viewing angle to a second viewing angle to obtain a pseudo label of the first sample at the second viewing angle includes: Compare the first sample with the moving object feature under the second viewing angle to obtain a third similarity vector mapped from the first sample to the second viewing angle; Substituting the moving object features under the second viewing angle into the softmax function for calculation, to obtain a first relationship matrix between target objects under the second viewing angle; adjusting the third similarity vector according to the first relationship matrix to obtain a fourth similarity vector; Determine effective moving object features under the second viewing angle, and adjust the fourth similarity vector according to the effective moving object features to obtain the pseudo label.
5. The method according to claim 4, characterized in that Determining effective moving object features under the second viewing angle includes: For any moving object feature under the second viewing angle, compare the moving object feature with all moving object features under the first viewing angle to obtain a likelihood vector; The likelihood vector is compared with a preset threshold, and if the likelihood vector is greater than the preset threshold, the motion object feature corresponding to the likelihood vector is determined to be the valid motion object feature.
6. The method according to claim 4, characterized in that The step of self-mapping the first sample to the first viewing angle to obtain a soft label of the first sample at the first viewing angle includes: For any moving object feature under the first viewing angle, compare the moving object feature with all moving object features under the second viewing angle to obtain a likelihood vector; Comparing the likelihood vector with a preset threshold, and if the likelihood vector is greater than the preset threshold, determining that the moving object feature corresponding to the likelihood vector is a valid moving object feature; The fourth similarity vector is adjusted according to the effective moving object feature to obtain the soft label.
7. The method according to claim 3, characterized in that Determining the relationship between the sample groups according to the pseudo labels and the soft labels includes: Selecting at least two samples from the first viewing angle to form the sample group, wherein the at least two samples include the first sample; When the sample group includes two samples, determining a binary relationship between the sample groups according to the pseudo label and the soft label, wherein the binary relationship is used to indicate the similarity between the two samples, wherein the binary relationship includes one of the following relationships: KL divergence, Euclidean distance, and JS divergence; When the sample group includes three samples, a triplet relationship between the sample groups is determined according to the pseudo label and the soft label, and the triplet relationship is used to indicate the similarity between the three samples.
8. The method according to claim 3, characterized in that Determining a second loss function when transferring the moving object features under the multiple viewing angles according to the relationship between the sample groups includes: Substituting the relationship between the sample groups into a smooth L1 loss function or a mean square error loss function for calculation, to obtain a loss function for transferring moving object features between the first perspective and the second perspective; The loss function when the moving object feature is transferred between any two perspectives under all perspectives is calculated, and all the obtained loss functions are accumulated and summed to obtain the second loss function.
9. The method according to claim 1, characterized in that: Determining the relationship between the features of the moving object under different viewing angles according to the first loss function, the second loss function, and the third loss function includes: Taking a weighted sum of the first loss function, the second loss function and the third loss function to obtain a fourth loss function; The relationship between the features of the moving object under different perspectives is determined based on the fourth loss function.
10. A moving object re-identification method, characterized in that: include: Acquire a target moving object image to be identified; Inputting the target moving object image into a moving object re-recognition model for analysis to obtain an image recognition result; The moving object re-identification model is obtained by obtaining sample subsets of moving objects under multiple perspectives, each sample subset contains moving object features of all moving objects under one perspective, the moving objects are objects that appear under the multiple perspectives during movement, each of the sample subsets is input into the target learning model for analysis, and the moving object features under each perspective are obtained, and the relationship between the moving object features under different perspectives is determined based on the moving object features under each perspective, and the target learning model is trained based on the relationship between the moving object features under each perspective and the moving object features under different perspectives; Among them, according to the motion object features at each viewing angle, the relationship between the motion object features at different viewing angles is determined, including: according to the motion object features corresponding to each motion object sample in each sample subset, the loss function at the viewing angle corresponding to each sample subset is determined, and the loss functions at multiple viewing angles are accumulated and summed to obtain a first loss function; according to the motion object features at each viewing angle, a second loss function is determined when the motion object features are transmitted at the multiple viewing angles; for any viewing angle, a first mean vector and a first standard deviation vector of all motion object features at the viewing angle are calculated, and a second mean vector and a second standard deviation vector of all motion object features at all viewing angles are calculated, and according to the first mean vector, the first standard deviation vector, the second mean vector and the second standard deviation vector, a third loss function is determined when the motion object feature distributions at the multiple viewing angles are aligned; according to the first loss function, the second loss function and the third loss function, the relationship between the motion object features at different viewing angles is determined.
11. The method according to claim 10, characterized in that The target moving object image is input into a moving object re-identification model for analysis to obtain an image recognition result, including: Determine an application mode of the current moving object re-identification model; the application mode includes a moving object authentication mode and a moving object retrieval mode, wherein the moving object authentication mode is used to determine whether the moving objects in the plurality of target moving object images are the same moving object, and the moving object retrieval mode is used to retrieve all images of the moving object corresponding to the target moving object image from a database, wherein the database includes images of a plurality of moving objects under a plurality of viewing angles; When it is determined that the current application mode is the moving object authentication mode, the moving object re-identification model analyzes the moving object features corresponding to the plurality of target moving object images, determines the similarity between the plurality of moving object features and compares them with a preset threshold, and determines whether the moving objects in the plurality of target moving object images are the same moving object based on the comparison result; When it is determined that the current application mode is the motion object retrieval mode, the motion object re-identification model analyzes the motion object features corresponding to the target motion object image, determines the similarity between the motion object features and the motion object features corresponding to all motion object images in the database, and compares them with a preset threshold, and determines all images of the motion object corresponding to the target motion object image in the database based on the comparison result.
12. A model training device, characterized in that: include: An acquisition module, used to acquire sample subsets of moving objects under multiple viewing angles, wherein each sample subset contains moving object features of all moving objects under one viewing angle, and the moving objects are objects that appear under the multiple viewing angles during movement; An analysis module, used for inputting each of the sample subsets into a target learning model for analysis to obtain the characteristics of the moving object at each viewing angle; A determination module, used to determine the relationship between the moving object features under different viewing angles according to the moving object features under each viewing angle, including: determining the loss function under the viewing angle corresponding to each sample subset according to the moving object features corresponding to each moving object sample in each sample subset, and accumulating and summing the loss functions under multiple viewing angles to obtain a first loss function; determining a second loss function when transferring the moving object features under the multiple viewing angles according to the moving object features under each viewing angle; for any one viewing angle, calculating the first mean vector and the first standard deviation vector of all the moving object features under the viewing angle, calculating the second mean vector and the second standard deviation vector of all the moving object features under all viewing angles, and determining a third loss function when aligning the distribution of the moving object features under the multiple viewing angles according to the first mean vector, the first standard deviation vector, the second mean vector and the second standard deviation vector; determining the relationship between the moving object features under the different viewing angles according to the first loss function, the second loss function and the third loss function; The training module is used to train the target learning model according to the relationship between the moving object features under each viewing angle and the moving object features under different viewing angles to obtain a moving object re-identification model.
13. A non-volatile storage medium, wherein: The non-volatile storage medium includes a stored program, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the model training method described in any one of claims 1 to 11.
14. An electronic device, comprising: processor; as well as A memory is connected to the processor and is used to provide the processor with instructions for processing the following processing steps: obtaining sample subsets of moving objects under multiple perspectives, wherein each sample subset contains moving object features of all moving objects under one perspective, and the moving objects are objects that appear under the multiple perspectives during movement; inputting each of the sample subsets into a target learning model for analysis to obtain moving object features under each perspective; determining the relationship between moving object features under different perspectives based on the moving object features under each perspective; training the target learning model based on the relationship between the moving object features under each perspective and the moving object features under different perspectives to obtain a moving object re-identification model; Among them, according to the motion object features at each viewing angle, the relationship between the motion object features at different viewing angles is determined, including: according to the motion object features corresponding to each motion object sample in each sample subset, the loss function at the viewing angle corresponding to each sample subset is determined, and the loss functions at multiple viewing angles are accumulated and summed to obtain a first loss function; according to the motion object features at each viewing angle, a second loss function is determined when the motion object features are transmitted at the multiple viewing angles; for any viewing angle, a first mean vector and a first standard deviation vector of all motion object features at the viewing angle are calculated, and a second mean vector and a second standard deviation vector of all motion object features at all viewing angles are calculated, and according to the first mean vector, the first standard deviation vector, the second mean vector and the second standard deviation vector, a third loss function is determined when the motion object feature distributions at the multiple viewing angles are aligned; according to the first loss function, the second loss function and the third loss function, the relationship between the motion object features at different viewing angles is determined.
Citation Information
Patent Citations
Cross-camera re-identification fusion method and system for similar-appearance targets
CN109800794A