Model Training Method, Device, and Moving Object Re-identification Method

In the training of the motion object re-recognition model, a subset of training data from multiple perspectives is obtained and divided into data sets of different categories for iterative training, the problem of low recognition accuracy caused by the difference between the training scenario and the actual application scenario is solved, and a higher cross-scene recognition accuracy is achieved.

CN115147452BActive Publication Date: 2025-05-30ALIBABA INNOVATION PRIVATE LIMITED
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110342267.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-30
Publication Date
2025-05-30
Estimated Expiration
2041-03-30

AI Technical Summary

Technical Problem

There are differences between the training scenarios and actual application scenarios in the motion object re-recognition model, resulting in low recognition accuracy.

Method used

By obtaining multiple training data subsets corresponding to multiple perspectives, dividing them into the first type of data set and the second type of data set, and inputting the target learning model in each iteration step for training, optimizing the model parameters.

Benefits of technology

The recognition accuracy of the moving object re-recognition model in cross-scene is improved, and the accuracy reduction caused by the difference between training scenarios and actual application scenarios is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115147452B_ABST
    Figure CN115147452B_ABST
Patent Text Reader

Abstract

The present application discloses a model training method, an apparatus, and a moving object re-identification method. Among them, the model training method includes: obtaining a plurality of training data subsets corresponding to a plurality of perspectives, each training data subset including all moving object images under one perspective, and the moving object is an object that appears under multiple perspectives during movement; dividing the plurality of training data subsets into a first type of data set and a second type of data set; in each iteration step, inputting the first type of data set into the target learning model, training the model based on the initial model parameters to determine the first loss function, and determining the update amount of the initial model parameters; inputting the second type of data set into the target learning model, training the model based on the update amount of the initial model parameters to determine the second loss function, and determining the target model parameters; and iteratively updating the target learning model according to the target model parameters determined in each iteration step.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of machine learning, and more particularly, to a model training method, apparatus, and moving object re-identification method. Background Art

[0002] The moving object re-identification technology has a large number of applications in many fields such as security, retail, and multimedia content management. In order to improve the accuracy of the moving object re-identification model, it is necessary to collect and label a high-quality training data set, that is, the data set contains a large number of moving objects with different identities, each moving object contains multiple images, and these images are from different shooting environments.

[0003] The most important challenge in the actual application of the moving object re-identification technology is that there are differences between the acquisition scenarios of the training data set and the scenarios where the moving object re-identification model is used, resulting in a decrease in the model accuracy. For example, the training data set is collected in an indoor environment, while the actual application scenario is an outdoor environment. Due to the different lighting conditions indoors and outdoors, the appearance of the moving object images changes, and the accuracy of the moving object re-identification model will be significantly reduced. Another example is that the training data set is collected in a small-scale scenario, while the actual application scenario is a larger-scale scenario. Since the installation height of the shooting lens in the latter is higher, the resolution of the moving object in the image is lower and the detailed information is less. If the moving object re-identification model trained on the data set in the former scenario is directly used in the latter scenario, the accuracy will also be significantly reduced. In order to obtain the highest possible moving object re-identification accuracy, it is necessary to solve the problem of differences between the training scenario and the actual application scenario. These differences include a large number of factors such as imaging perspective, imaging distance, background in the moving object image, lens imaging style, and lighting conditions, which cannot be solved by simple image processing means and must be improved in terms of training data and model training methods.

[0004] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention

[0005] Embodiments of the present application provide a model training method, apparatus, and moving object re-identification method to at least solve the technical problem of low recognition accuracy in moving object re-identification due to differences between the training scenario and the actual application scenario.

[0006] According to one aspect of the embodiments of the present application, there is provided a model training method, including: obtaining a plurality of training data subsets corresponding to a plurality of perspectives, where each training data subset includes all moving object images under one perspective, and the moving object is an object that appears under the plurality of perspectives during movement; dividing the plurality of training data subsets into a first type of data set and a second type of data set; in each iteration step, inputting the first type of data set into a target learning model to train the target learning model based on a first loss function determined by initial model parameters, and determining an update amount of the initial model parameters; inputting the second type of data set into the target learning model, and training the target learning model with a second loss function determined based on the update amount of the initial model parameters to determine target model parameters; and iteratively updating the target learning model according to the target model parameters determined in each iteration step.

[0007] According to another aspect of the embodiments of the present application, there is also provided another model training method, including: obtaining a training data set; dividing the training data set into a first type of data set and a second type of data set; inputting the first type of data set into a target learning model for training to determine an update amount of initial model parameters; inputting the second type of data set into the target learning model for training to determine target model parameters; and updating the target learning model based on the target model parameters.

[0008] According to another aspect of the embodiments of the present application, there is also provided a moving object re-identification method, including: obtaining a target moving object image to be identified; inputting the target moving object image into a moving object re-identification model for analysis to obtain an image recognition result; where the moving object re-identification model is trained in the following manner: obtaining a plurality of training data subsets corresponding to a plurality of perspectives, where each training data subset includes all moving object images under one perspective, and the moving object is an object that appears under the plurality of perspectives during movement; dividing the plurality of training data subsets into a first type of data set and a second type of data set; in each iteration step, inputting the first type of data set into a target learning model to train the target learning model based on a first loss function determined by initial model parameters, and determining an update amount of the initial model parameters; inputting the second type of data set into the target learning model, and training the target learning model with a second loss function determined based on the update amount of the initial model parameters to determine target model parameters; and iteratively updating the target learning model according to the target model parameters determined in each iteration step.

[0009] According to another aspect of the embodiments of the present application, there is also provided a model training device, including: an acquisition module, configured to acquire a plurality of training data subsets corresponding to a plurality of perspectives, where each training data subset includes all moving object images under one perspective, and the moving object is an object that appears under the plurality of perspectives during movement; a classification module, configured to divide the plurality of training data subsets into a first type of data set and a second type of data set; a determination module, configured to, in each iteration step, input the first type of data set into a target learning model to train the target learning model based on a first loss function determined by initial model parameters, and determine an update amount of the initial model parameters; input the second type of data set into the target learning model, and train the target learning model based on a second loss function determined by the update amount of the initial model parameters to determine target model parameters; an update module, configured to iteratively update the target learning model according to the target model parameters determined in each iteration step.

[0010] According to another aspect of the embodiments of the present application, there is also provided a non-volatile storage medium, where the non-volatile storage medium includes a stored program, and when the program runs, it controls the device where the non-volatile storage medium is located to execute the above-mentioned model training method.

[0011] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including: a processor and a memory connected to the processor, where the memory is used to provide instructions for the processor to perform the following processing steps: acquire a plurality of training data subsets corresponding to a plurality of perspectives, where each training data subset includes all moving object images under one perspective, and the moving object is an object that appears under the plurality of perspectives during movement; divide the plurality of training data subsets into a first type of data set and a second type of data set; in each iteration step, input the first type of data set into a target learning model to train the target learning model based on a first loss function determined by initial model parameters, and determine an update amount of the initial model parameters; input the second type of data set into the target learning model, and train the target learning model based on a second loss function determined by the update amount of the initial model parameters to determine target model parameters; iteratively update the target learning model according to the target model parameters determined in each iteration step.

[0012] In the embodiments of the present application, by dividing the training data set into a first type of data set and a second type of data set with non-overlapping perspectives, and inputting them into the target learning model for training respectively, the model parameters are jointly optimized and adjusted, which is similar to repeatedly performing the training and testing processes. The image data used in the testing process is different from that in the training process, thereby solving to a certain extent the problem of accuracy decline caused by the inconsistency between the training data set and the target scenario. At the same time, the present application further improves the accuracy of re-identifying moving objects across scenarios by introducing cross-perspective loss, so that the obtained moving object re-identification model can still have a high accuracy when deployed to the target scenario without obtaining the target scenario data, thus solving the technical problem of low recognition accuracy in moving object re-identification due to the difference between the training scenario and the actual application scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:

[0014] Figure 1 is a schematic hardware structure diagram of a computer terminal according to Embodiment 1 of the present application;

[0015] Figure 2 is a schematic flowchart of a model training method according to Embodiment 1 of the present application;

[0016] Figure 3 is a schematic diagram of a training data set of a moving object re-identification model according to Embodiment 1 of the present application;

[0017] Figure 4 is a schematic flowchart of a model training method according to Embodiment 2 of the present application;

[0018] Figure 5 is a schematic flowchart of a moving object re-identification method according to Embodiment 3 of the present application;

[0019] Figure 6a is an application schematic diagram of a pedestrian re-identification method according to Embodiment 3 of the present application;

[0020] Figure 6b is an application schematic diagram of another pedestrian re-identification method according to Embodiment 3 of the present application;

[0021] Figure 7 is a schematic structural diagram of a model training device according to Embodiment 4 of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solution in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0023] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0024] First, some nouns or terms that appear in the process of describing the embodiments of this application are applicable to the following explanations:

[0025] Re-identification of moving objects: Also known as re-identification of moving objects, it is a technology that uses computer vision technology to determine whether there are specific moving objects in an image or video sequence. It is widely regarded as a sub-problem of image retrieval. The re-identification of moving objects aims to learn a model that can extract identity features for images of moving objects, that is, extract features that can fully represent the appearance of the moving object. Based on this model, images of the same moving object under different cameras can be matched together. By given an image of a moving object, the images of this moving object across devices can be retrieved, which can make up for the visual limitations of fixed cameras and can be combined with moving object detection / tracking technology, and can be widely applied in fields such as intelligent security.

[0026] Training dataset for re-identification of moving objects (abbreviation: training set): The dataset contains a large number of images of moving objects with different identities (IDs). Each ID contains multiple images, and each image has an ID label. Adding an identity ID label to an image without an ID label is called data or image annotation, abbreviated as annotation.

[0027] Moving object re-identification feature (abbreviated as feature): In the moving object re-identification task, a moving object image is represented by a digital vector, and this digital vector is the moving object re-identification feature of this image. Generally speaking, for two moving object images with the same appearance, the similarity of their moving object re-identification vectors is relatively high; for two moving object re-identification images with different appearances, the similarity of their moving object re-identification vectors is relatively low.

[0028] Moving object re-identification model (abbreviated as model): Used to obtain the moving object re-identification feature from a moving object image.

[0029] Embodiment 1

[0030] In the related art, to solve the problem that there are differences between the training scenario and the target scenario, one solution is to collect moving object image data in the target scenario and perform annotation. Using these annotated data, the model that has been trained in the training scenario is further optimized. The principle of this solution is simple. Because after annotation, it is equivalent to generating a new training data set collected from the target scenario. This data set can be directly used to train the existing moving object re-identification model, or the new data set can be merged with the original data set to train the moving object re-identification model, and the obtained moving object re-identification model can adapt to the target scenario. However, its disadvantage is that the cost of annotating moving object images is very high, because only after collecting and annotating a large number of training samples in the target scenario can the model better adapt to the target scenario.

[0031] Another solution is to collect images of moving objects in the target scenario without annotation, but directly use these unannotated images of moving objects to optimize the model in the target scenario. Such methods are generally called unsupervised domain adaptation methods, and there are two common implementation means. The first means is to use the images of moving objects collected in the target scenario as a reference, perform style transfer on the images of moving objects in the training set of the original moving object re-identification model, so that the images in the original training set have a similar appearance to the images in the target scenario, and then use these training samples that have been style transferred and already have labels to train the model. In this way, the obtained model can adapt to the target scenario to a certain extent. The second means is to use an existing moving object re-identification model to extract the moving object re-identification features of the images of moving objects collected in the unannotated target scenario, and then based on these features, cluster the images of moving objects in the target scenario, and assign an ID label to each image of moving objects in the target scenario according to the clustering result, that is, use the existing moving object re-identification model to perform automatic image annotation on the images of moving objects in the target scenario, and then use the annotated images of moving objects in the target scenario, or use the annotated images of moving objects in the target scenario and the images in the original training set together to train the moving object re-identification model. Its disadvantage is that it still needs to collect images of moving objects in the target scenario. When the number of target scenarios is large, it takes a large cost to collect images of each target scenario; in addition, for commercial applications, the images of the target scenario usually belong to the privacy information of customers, and it is very likely that it is impossible to obtain a large number of customer images of the target scenario for model optimization.

[0032] To solve the above problems, the embodiment of the present application proposes a model training method based on domain generalization, which does not need to collect and annotate images of moving objects in the target scenario, but through the optimized design of the training process, enables the model to have stronger generalization ability, thereby improving the accuracy of moving object re-identification in various target scenarios. The following is a detailed description.

[0033] It should be noted that the moving objects in the embodiment of the present application are not limited to pedestrians, but can also be other objects that appear in multiple shots in sequence, such as pets on the street, animals in the zoo, etc.

[0034] The method embodiment provided by the embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 The hardware structure block diagram of a computer terminal (or mobile device) for implementing the model training method is shown. As Figure 1As shown, the computer terminal 10 (or mobile device 10) may include one or more processors 102 (shown as 102a, 102b, ……, 102n in the figure) (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than those Figure 1 shown in, or have a different configuration from that Figure 1 shown.

[0035] It should be noted that the above one or more processors 102 and / or other data processing circuits can generally be referred to as "data processing circuits" in this article. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 10 (or mobile device). As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistor terminal path connected to an interface).

[0036] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the model training method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the vulnerability detection method of the above-mentioned application program. The memory 104 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely set relative to the processor 102, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0037] The transmission module 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission module 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission module 106 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0038] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables a user to interact with the user interface of the computer terminal 10 (or mobile device).

[0039] Under the above operating environment, an embodiment of the present application provides a model training method for re-identifying moving objects. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0040] As Figure 2 shown, the specific process of the model training method includes steps S202 - S208, where:

[0041] Step S202, obtaining multiple training data subsets corresponding to multiple perspectives, where each training data subset includes all moving object images under one perspective, and the moving object is an object that appears in multiple perspectives in sequence during movement. That is to say, the training data subsets are obtained by dividing the initial training data set according to perspectives, where one training data subset corresponds to one perspective.

[0042] In some optional embodiments of the present application, multiple moving object images under multiple perspectives can be obtained, and all moving object images under each perspective are used as a training data subset. Among them, the field of view range of one camera can be used as one perspective, that is, one camera corresponds to one perspective; of course, different shooting angles of one camera can also be used as different perspectives, that is, one camera can correspond to multiple perspectives.

[0043] For example, images of moving objects captured by multiple cameras set at different locations are collected as a training data set, and the images of moving objects are divided into multiple training data subsets according to the cameras corresponding to these images of moving objects. Each training data subset includes all the images of moving objects captured by one camera. Since the cameras come from multiple different locations, there are many variations in the backgrounds, lighting conditions, perspectives, and poses of the moving objects in the captured images of moving objects. Using such a training data set, through a deep learning algorithm, the obtained moving object re-identification model will have better robustness to scene changes.

[0044] Each obtained image of a moving object has two labels. One is an identity label (ID), which is used to indicate the identity of the moving object corresponding to the image of the moving object; the other is a perspective label (Cam), which is used to indicate the perspective to which the image of the moving object belongs, that is, the camera from which it is sourced. An optional training data set is as Figure 3 shown. Classify the images of moving objects according to the label information of each image of a moving object. Each cell includes all the images of a moving object under one perspective. On the one hand, each row is a set of all the images of the same moving object, and its corresponding perspective labels include all the perspectives in the training data set; on the other hand, each column is a set of all the images of moving objects under the same perspective, that is, the training data subset in the embodiment of the present application, and its corresponding identity labels include all the moving objects in the training data set.

[0045] It should be noted that in practice, on the one hand, there are some moving object IDs that only come from some perspectives and cannot cover all perspectives; on the other hand, there are some perspectives where the obtained images of moving objects only contain some of the moving object IDs rather than all of them; that is Figure 3 the number of images in some cells in the table is 0. However, this situation does not affect the actual use of the model training method in the embodiment of the present application, because the training of the moving object re-identification model is a process that includes a large number of iterative steps. In each single step of each iteration, only a small number of images of moving objects are needed to optimize the parameters of the model, as long as it is ensured that the corresponding training sample groups can be selected according to the preset rules in each single step of each iteration.

[0046] Step S204, divide the multiple training data subsets into a first type of data set and a second type of data set.

[0047] In some optional embodiments of the present application, during any training cycle of the training process, the multiple training data subsets need to be randomly divided into two parts to obtain a first type of data set (also called a support data set) and a second type of data set (also called a query data set) respectively. Taking Figure 3Taking the training data set in as an example, the moving object images corresponding to perspective 1, perspective 3, and perspective 4 are combined to form a support data set, and the moving object images corresponding to perspective 2 and perspective 5 are combined to form a query data set, that is, it is ensured that there is no perspective overlap in the moving object images in the support data set and the query data set.

[0048] Among them, the training cycle refers to the process of traversing all the data in a complete training data set. The training process of the actual moving object re-identification model needs to traverse the training data set multiple times, that is, it will include multiple training cycles; within each training cycle, the model training also includes multiple iteration processes, and each iteration process is called a training step or an iteration step.

[0049] Step S206, in each iteration step, input the first type of data set into the target learning model, train the target learning model with the first loss function determined based on the initial model parameters, and determine the update amount of the initial model parameters; input the second type of data set into the target learning model, and train the target learning model with the second loss function determined based on the update amount of the initial model parameters to determine the target model parameters.

[0050] In some optional embodiments of the present application, by inputting the support data set and the query data set into the target learning model for training respectively, the update amount of the initial model parameters and the final target model parameters are obtained in sequence, and the moving object re-identification model is updated according to the target model parameters. Specifically, this process includes the following steps S2061-S2066, where:

[0051] Step S2061, select the first training sample group from the first type of data set based on a preset rule.

[0052] Step S2062, select the second training sample group from the second type of data set based on a preset rule.

[0053] In some optional embodiments of the present application, determine the first preset number of identity labels from the first type of data set, and select the second preset number of moving object images from the multiple moving object images corresponding to each identity label to form the first training sample group; determine the first preset number of identity labels from the second type of data set, and select the second preset number of moving object images from the multiple moving object images corresponding to each identity label to form the second training sample group; among them, the first preset number of identity labels determined in the second type of data set and the first type of data set are exactly the same.

[0054] For example, in each iterative single step, 16 moving object IDs are randomly selected from the support dataset. For each ID, 4 moving object images are selected. The 64 obtained moving object images form the first training sample group, and the moving object images therein are called support images. Then, the same 16 moving object IDs are selected from the query dataset. For each ID, 4 moving object images are selected. The 64 obtained moving object images form the second training sample group, and the moving object images therein are called query images. Here, whether it is the first training sample group or the second training sample group, the 4 moving object images corresponding to each ID may come from different perspectives or may all come from the same perspective. It should be noted that the numbers 16, 4, and 64 are only used for illustration here and are not specifically limited. In actual training, different parameters can be selected according to the size of the training dataset.

[0055] Step S2063: Input the first training sample group into the target learning model, and determine the first loss function based on the initial model parameters of the target learning model.

[0056] Step S2064: Determine the update amount of the initial model parameters based on the first loss function.

[0057] In some optional embodiments of the present application, based on the moving object images in the first training sample group and the initial model parameters of the target learning model, the first loss function is calculated. The first loss function includes at least one of the following loss functions: cross-entropy loss function, contrastive loss function, triplet loss function, center loss function.

[0058] Specifically, first input the support images in the first training sample group into the target training model, and determine the first loss function (also called the basic loss function). The update amount of the initial model parameters is obtained through the basic loss function. The basic loss function can adopt forms such as cross-entropy loss, contrastive loss, triplet loss, and center loss commonly used in the training of moving object re-identification models. One of them can be selected, or multiple losses can be combined for calculation. Although this step is similar to the usual moving object re-identification algorithm, there is an important difference: after obtaining the update amount of the initial model parameters, it is not used to update the model parameters but is used for subsequent calculations.

[0059] An optional basic loss function provided by the embodiments of the present application is as follows. Here, θ is the initial model parameter of the target learning model, D s represents the support image, and L B (θ) represents the total loss of the basic loss function, which is composed of the cross-entropy loss function L soft (D s ; θ) and the triplet loss function L tri (D s; θ) is composed of two parts:

[0060] L B (θ) = L soft (D s ; θ) + L tri (D s ; θ)

[0061] According to the basic loss function L B (θ), the update amount of the initial model parameters is obtained where α is a preset update step size, denotes the partial derivative with respect to θ.

[0062] Step S2065: Input the second training sample group into the target learning model, and determine the second loss function based on the update amount of the initial model parameters.

[0063] Step S2066: Determine the target model parameters based on the second loss function.

[0064] In some alternative embodiments of the present application, the second model parameters are determined based on the initial model parameters and the update amount of the initial model parameters; based on the moving object images in the second training sample group and the second model parameters, the third loss function is calculated, and the third loss function has the same structure as the first loss function, and its parameters are the second model parameters; based on the moving object images in the first training sample group, the moving object images in the second training sample group, and the second model parameters, the cross-view loss function is calculated; the third loss function and the cross-view loss function are summed to obtain the second loss function.

[0065] Specifically, first determine the second model parameters θ′ according to the update amount of the initial model parameters determined in the previous step. It should be noted that θ′ is only used for subsequent calculations here and does not update the model parameters of the target learning model, where

[0066]

[0067] After that, input the query images in the second training sample group into the target training model, and determine the second loss function (also known as the generalization loss function). The target model parameters are obtained by minimizing the generalization loss function, and the target model parameters are used to update the model parameters of the target learning model. Generally, the calculation formula of the generalization loss function is as follows:

[0068] L G (θ′) = L soft (D q ; θ′) + L tri (D q ; θ′) + L ccm ((D s , D q); θ')

[0069] Among them, D q represents the query image. It can be seen that the generalization loss function includes two parts. The first part is the third loss function, that is, L soft (D q ; θ') + L tri (D q ; θ'). Its structure is the same as that of the basic loss function, and the only difference is that the parameter in the third loss function is the second model parameter θ'; the second part is the cross-view loss function L ccm ((D s , D q ); θ').

[0070] In some optional embodiments of the present application, the process of determining the cross-view loss function includes: analyzing the motion object features corresponding to all the motion object images in the first training sample group and the second training sample group, and calculating the mean value of all the motion object features in the first training sample group; calculating the similarity between each motion object feature in the second training sample group and the mean value, and determining the cross-view loss function based on the similarity. An optional expression formula for the cross-view loss function is:

[0071]

[0072] Among them, x q represents the q-th query image, and f θ′ (x q ) is the motion object feature of this query image, is the mean value of the motion object features of all the support images with the p-th ID. In the numerator, the query image and the support image for calculating the cosine similarity have the same motion object ID; in the denominator, the similarity between the motion object feature from the query image and the mean value of the motion object features from the support image is calculated for all the motion object IDs, which is used to normalize the numerator. Since the optimization is essentially carried out using the query image, so and from the support image can be considered as constants rather than functions of θ', and they do not participate in the chain rule of the composite function derivative, which can simplify the calculation process.

[0073] After determining the generalization loss function, by adjusting the initial model parameter θ to minimize the generalization loss function, the update amount of the initial model parameter θ is further determined to obtain the final target model parameter:

[0074]

[0075] Among them, β is a preset update step size. When calculating the partial derivative of θ, according to the derivative rule of composite functions, it is necessary to first calculate the partial derivative of θ′, which is obtained by inputting the query image into the model with θ′ as the parameter; then calculate the partial derivative of θ′ with respect to θ, which requires calculating the second-order derivative of the basic loss function with respect to θ, obtained by inputting the support image into the model with θ as the parameter, and finally multiply these two terms.

[0076] Step S208: Iteratively update the target learning model according to the target model parameters determined in each iterative single step.

[0077] Through the above process, at the end of each iterative single step, a target model parameter can be obtained, which is used to update the model parameters of the target learning model. Since a training cycle includes multiple iterative single steps, and the entire training process includes multiple training cycles, the training process of the moving object re-identification model is essentially to repeatedly train the target learning model through the support image and the query image, iteratively update and optimize the model parameters, and finally obtain a reliable moving object re-identification model.

[0078] In the embodiments of the present application, by dividing the training data set into a first type of data set and a second type of data set with non-overlapping perspectives and inputting them into the target learning model for training respectively, jointly optimizing and adjusting the model parameters, which is similar to repeatedly performing the training and testing processes. The image data used in the testing process is different from that in the training process, thus solving to a certain extent the problem of accuracy decline caused by the inconsistency between the training data set and the target scenario; at the same time, the present application further improves the accuracy of cross-scenario moving object re-identification by introducing cross-perspective loss, so that the obtained moving object re-identification model can still have a high accuracy when deployed to the target scenario without obtaining the target scenario data, thereby solving the technical problem of low recognition accuracy in moving object re-identification due to the difference between the training scenario and the actual application scenario.

[0079] Embodiment 2

[0080] According to the embodiments of the present application, another embodiment of the model training method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0081] The model training method provided by the embodiments of the present application can also run in Figure 1 the operating environment shown, and its flowchart is as Figure 4 shown. This method includes steps S402 - S410, where:

[0082] Step S402: Obtain a training data set.

[0083] In some optional embodiments of the present application, multiple moving object images from multiple perspectives can be obtained, and all the moving object images from each perspective are used as a training data subset. The moving object is an object that appears in multiple perspectives during movement. For example, moving object images captured by multiple cameras at different locations are used as the training data set, and they are divided into multiple training data subsets according to the cameras corresponding to these moving object images. Each training data subset includes all the moving object images captured by one camera. Since the cameras are from multiple different locations, the moving object images captured are necessarily of different styles. Therefore, the moving object re-identification model trained using the obtained training data set can have better robustness to the environment. Each obtained moving object image has two labels. One is an identity label (ID), which is used to indicate the identity of the moving object corresponding to the moving object image. The other is a perspective label (Cam), which is used to indicate the perspective to which the moving object image belongs, that is, the camera from which it is sourced. An optional training data set is as Figure 3 shown.

[0084] Step S404: Divide the training data set into a first type of data set and a second type of data set.

[0085] In some optional embodiments of the present application, during any training cycle of the training process, multiple training data subsets are randomly divided into two parts to obtain a first type of data set (also known as a support data set) and a second type of data set (also known as a query data set). Taking the Figure 3 training data set as an example, the moving object images corresponding to perspective 1, perspective 3, and perspective 4 are used to form the support data set, and the moving object images corresponding to perspective 2 and perspective 5 are used to form the query data set, that is, it is ensured that there is no perspective overlap between the moving object images in the support data set and the query data set.

[0086] Step S406: Input the first type of data set into the target learning model for training to determine the update amount of the initial model parameters.

[0087] In some alternative embodiments of the present application, a first preset number of identity tags are determined from a first type of dataset, and a second preset number of moving object images are selected from multiple moving object images corresponding to each identity tag to form a first training sample group; the support images in the first training sample group are input into a target training model, and a basic loss function is determined, and an update amount of the initial model parameters is obtained through the basic loss function. Among them, the basic loss function can adopt forms such as cross-entropy loss, contrastive loss, triplet loss, center loss, etc. commonly used in the training of moving object re-identification models. One of them can be selected, or multiple losses can be combined for calculation. This step is similar to the usual moving object re-identification algorithm, but there is an important difference: after obtaining the update amount of the initial model parameters, it is not used to update the model parameters, but is used for calculation in subsequent steps.

[0088] Step S408, input the second type of dataset into the target learning model for training to determine the target model parameters.

[0089] In some alternative embodiments of the present application, a first preset number of identity tags are determined from a second type of dataset, and a second preset number of moving object images are selected from multiple moving object images corresponding to each identity tag to form a second training sample group, where the first preset number of identity tags determined in the second type of dataset and the first type of dataset are exactly the same.

[0090] After that, first determine the second model parameters according to the update amount of the initial model parameters determined in the previous step, then input the query images in the second training sample group into the target training model, and determine the generalization loss function. The target model parameters are obtained by minimizing the generalization loss function, and the target model parameters are used to update the model parameters of the target learning model. Among them, the generalization loss function usually includes two parts. The first part is the basic loss function, whose structure is the same as the basic loss function in the previous step, and the only difference is that its parameters are the second model parameters; the second part is the cross-view loss function used to further optimize the model.

[0091] Step S410, update the target learning model based on the target model parameters.

[0092] Through the above process, in each iteration process, a target model parameter can be obtained, which is used to update the model parameters of the target learning model.

[0093] Since a training cycle includes multiple iterative steps, and the entire training process includes multiple training cycles, the training process of the moving object re-identification model is essentially to repeatedly train the target learning model through support images and query images, iteratively update and optimize the model parameters, and finally obtain a reliable moving object re-identification model.

[0094] In an embodiment of the present application, a training data set is obtained; the training data set is divided into a first type of data set and a second type of data set; the first type of data set is input into a target learning model for training to determine an update amount of initial model parameters; the second type of data set is input into the target learning model for training to determine target model parameters; the target learning model is updated based on the target model parameters; wherein, by dividing the training data set into a first type of data set and a second type of data set with non-overlapping perspectives and inputting them into the target learning model for training respectively, the model parameters are jointly optimized and adjusted, which to a certain extent solves the problem of accuracy decline caused by the inconsistency between the training data set and the target scenario; at the same time, by introducing a cross-perspective loss, the accuracy of re-identifying moving objects across scenarios is further improved, so that the obtained moving object re-identification model can still have a high recognition accuracy when deployed to the target scenario without obtaining the target scenario data.

[0095] Embodiment 3

[0096] According to an embodiment of the present application, an embodiment of a method for re-identifying moving objects is further provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0097] Figure 5 is a schematic flowchart of a method for re-identifying moving objects according to an embodiment of the present application. As Figure 5 shown, the method includes steps S502 - S504, where:

[0098] Step S502, obtain an image of a target moving object to be identified.

[0099] Step S504, input the image of the target moving object into a moving object re-identification model for analysis to obtain an image recognition result.

[0100] Among them, the moving object re-identification model is trained as follows: Obtain multiple training data subsets corresponding to multiple perspectives, where each training data subset includes all the moving object images under one perspective, and the moving object is an object that appears under multiple perspectives during movement; Divide the multiple training data subsets into a first type of data set and a second type of data set; In each iterative step, input the first type of data set into the target learning model, and train the target learning model based on the first loss function determined by the initial model parameters to determine the update amount of the initial model parameters; Input the second type of data set into the target learning model, and train the target learning model based on the second loss function determined by the update amount of the initial model parameters to determine the target model parameters; Iteratively update the target learning model according to the target model parameters determined in each iterative step.

[0101] Specifically, when training the moving object re-identification model, first obtain multiple moving object images under multiple perspectives, and use all the moving object images under each perspective as a training data subset; For example, collect the moving object images captured by multiple cameras at different locations as the training data set, and divide them into multiple training data subsets according to the cameras corresponding to these moving object images. Each training data subset includes all the moving object images captured by one camera.

[0102] After that, in any training cycle of the training process, randomly divide the multiple training data subsets into two parts to obtain a support data set and a query data set respectively; By inputting the support data set and the query data set into the target learning model for training respectively, obtain the update amount of the initial model parameters and the final target model parameters in sequence, and complete the update of the moving object re-identification model according to the target model parameters.

[0103] Since the moving object re-identification has two application modes: moving object authentication and moving object retrieval, when applying the moving object re-identification model, it is necessary to determine the application mode of the current moving object re-identification model; Among them, the moving object authentication mode is used to determine whether the moving objects in multiple target moving object images are the same moving object, and the moving object retrieval mode is used to retrieve all the images of the moving object in the target moving object image from the database, and the database includes the images of multiple moving objects under multiple perspectives.

[0104] When it is determined that the current application mode is the moving object authentication mode, the moving object re-identification model analyzes the moving object features corresponding to multiple target moving object images, determines the similarity between multiple moving object features and compares it with a preset threshold, and determines whether the moving objects in multiple target moving object images are the same moving object according to the comparison result; if the similarity between multiple moving object features is greater than the preset threshold, it can be confirmed that the moving objects in multiple target moving object images are the same moving object. Taking pedestrian re-identification as an example, an optional process of pedestrian re-identification is as Figure 6a shown, where camera 1 captures pedestrian image 1, and camera 2 captures pedestrian image 2. The two cameras respectively send the captured pedestrian images to the pedestrian re-identification model running on the server. The pedestrian re-identification model analyzes the pedestrian features corresponding to the two pedestrian images and makes a comparison, and finally determines whether the pedestrians captured by camera 1 and camera 2 are the same pedestrian.

[0105] When it is determined that the current application mode is the moving object retrieval mode, the moving object re-identification model analyzes the moving object features corresponding to the target moving object image, determines the similarity between the moving object features and the moving object features corresponding to all moving object images in the database, and compares it with a preset threshold. According to the comparison result, it determines all the images of the moving object corresponding to the target moving object image in the database. Specifically, all moving object images in the database whose feature similarity with the target moving object image is greater than the preset threshold can be considered as the images of the moving object corresponding to the target moving object image. Taking pedestrian re-identification as an example, an optional process of pedestrian re-identification is as Figure 6b shown. After camera 1 captures a pedestrian image, it sends the pedestrian image to the pedestrian re-identification model running on the server. The pedestrian re-identification model analyzes the pedestrian features corresponding to the pedestrian image, retrieves the database (which stores multiple pedestrian images captured by multiple cameras), compares the pedestrian image with all the pedestrian images in the database, and finally outputs the images of the pedestrian under all cameras. The pedestrian re-identification technology can be widely applied to fields such as shopping malls, transportation, exhibitions, and security.

[0106] It should be noted that the moving objects in the embodiments of this application are not limited to pedestrians, but can also be other objects that appear in multiple cameras, such as pets on the street, animals in the zoo, etc. For example, in the management of a wildlife park, given an animal image, the images of the animal in all cameras can be found through the above model to understand the activity range of the animal.

[0107] It should be noted that the moving object re-identification model in the embodiments of the present application is obtained according to the model training method in Embodiment 1. Since the training process has been described in detail in Embodiment 1, some details not shown in this embodiment can be referred to Embodiment 1 and will not be elaborated here.

[0108] Embodiment 4

[0109] According to the embodiments of the present application, there is also provided a model training device for implementing the above model training method, as Figure 7 shown. The device includes an acquisition module 70, a classification module 72, a determination module 74, and an update module 76, where:

[0110] The acquisition module 70 is configured to acquire multiple training data subsets corresponding to multiple perspectives, where each training data subset includes all moving object images under one perspective;

[0111] In some optional embodiments of the present application, multiple moving object images under multiple perspectives can be acquired, and all moving object images under each perspective are used as a training data subset. The moving object is an object that appears under multiple perspectives during movement. For example, moving object images captured by multiple cameras at different locations are collected as a training data set, and the moving object images are divided into multiple training data subsets according to the cameras corresponding to these moving object images. Each training data subset includes all moving object images captured by one camera. Each acquired moving object image has two labels, one is an identity label for indicating the identity of the moving object corresponding to the moving object image, and the other is a perspective label for indicating the perspective to which the moving object image belongs, that is, the camera from which it is sourced.

[0112] The classification module 72 is configured to divide the multiple training data subsets into a first type of data set and a second type of data set;

[0113] In some optional embodiments of the present application, during any training cycle of the training process, the multiple training data subsets are randomly divided into two parts to obtain a first type of data set (also referred to as a support data set) and a second type of data set (also referred to as a query data set) respectively.

[0114] The determination module 74 is configured to, in each iterative step, input the first type of data set into the target learning model to train the target learning model based on the first loss function determined by the initial model parameters, and determine the update amount of the initial model parameters; input the second type of data set into the target learning model, and train the target learning model based on the second loss function determined by the update amount of the initial model parameters to determine the target model parameters;

[0115] In some alternative embodiments of the present application, first, determine a first preset number of identity tags from a first type of dataset, and select a second preset number of moving object images from the multiple moving object images corresponding to each identity tag to form a first training sample group; determine a first preset number of identity tags from a second type of dataset, and select a second preset number of moving object images from the multiple moving object images corresponding to each identity tag to form a second training sample group; wherein, the first preset number of identity tags determined in the second type of dataset and the first type of dataset are exactly the same.

[0116] After that, input the support images in the first training sample group into the target training model, and determine a first loss function (also known as the basic loss function), and obtain the update amount of the initial model parameters through the basic loss function. Among them, the basic loss function can adopt forms such as cross-entropy loss, contrastive loss, triplet loss, center loss, etc. commonly used in the training of moving object re-identification models. One of them can be selected, or multiple losses can be combined for calculation.

[0117] Determine the second model parameters according to the determined update amount of the initial model parameters, then input the query images in the second training sample group into the target training model, and determine the generalization loss function. Obtain the target model parameters by minimizing the generalization loss function, and the target model parameters are used to update the model parameters of the target learning model. Among them, the generalization loss function generally includes two parts. The first part is the basic loss function, whose structure is the same as the basic loss function in the previous step, and the only difference is that its parameters are the second model parameters; the second part is the cross-view loss function used to further optimize the model.

[0118] The update module 76 is used to iteratively update the target learning model according to the target model parameters determined in each iterative single step.

[0119] Since a training cycle includes multiple iterative single steps, and the entire training process includes multiple training cycles, the training process of the moving object re-identification model is essentially to repeatedly train the target learning model through support images and query images, iteratively update and optimize the model parameters, and finally obtain a reliable moving object re-identification model.

[0120] It should be noted that each module in the model training device in the embodiments of the present application corresponds one by one to the implementation steps of the model training method in Embodiment 1. Since Embodiment 1 has been described in detail, some details not shown in this embodiment can be referred to Embodiment 1 and will not be elaborated here.

[0121] Embodiment 5

[0122] According to an embodiment of the present application, an electronic device is further provided. The electronic device includes a processor and a memory, where: The memory is connected to the processor and is used to provide instructions for the processor to process the following processing steps:

[0123] Obtain multiple training data subsets corresponding to multiple perspectives. Each training data subset includes all moving object images under one perspective, and the moving object is an object that appears under multiple perspectives during movement; divide the multiple training data subsets into a first type of data set and a second type of data set; in each iteration step, input the first type of data set into the target learning model to train the target learning model based on the first loss function determined by the initial model parameters, and determine the update amount of the initial model parameters; input the second type of data set into the target learning model and train the target learning model with the second loss function determined based on the update amount of the initial model parameters to determine the target model parameters; and iteratively update the target learning model according to the target model parameters determined in each iteration step.

[0124] Optionally, the memory further stores instructions for processing the following steps: Obtain a training data set; divide the training data set into a first type of data set and a second type of data set; input the first type of data set into the target learning model for training to determine the update amount of the initial model parameters; input the second type of data set into the target learning model for training to determine the target model parameters; and update the target learning model based on the target model parameters.

[0125] Optionally, the memory further stores instructions for processing the following steps: Obtain an image of a target moving object to be recognized; input the image of the target moving object into a moving object re-identification model for analysis to obtain an image recognition result.

[0126] Embodiment 6

[0127] According to an embodiment of the present application, a non-volatile storage medium is further provided. The non-volatile storage medium includes a stored program, where when the program runs, it controls the device where the non-volatile storage medium is located to execute the above-mentioned model training method.

[0128] Optionally, when the program is running, control the device where the non-volatile storage medium is located to perform the following steps: obtain multiple subsets of training data corresponding to multiple perspectives, where each subset of training data includes all the moving object images under one perspective, and the moving object is an object that appears under multiple perspectives during movement; divide the multiple subsets of training data into a first type of data set and a second type of data set; in each iteration step, input the first type of data set into the target learning model to train the target learning model based on the first loss function determined by the initial model parameters, and determine the update amount of the initial model parameters; input the second type of data set into the target learning model and train the target learning model with the second loss function determined based on the update amount of the initial model parameters to determine the target model parameters; and iteratively update the target learning model according to the target model parameters determined in each iteration step.

[0129] Optionally, when the program is running, control the device where the non-volatile storage medium is located to perform the following steps: obtain a training data set; divide the training data set into a first type of data set and a second type of data set; input the first type of data set into the target learning model for training to determine the update amount of the initial model parameters; input the second type of data set into the target learning model for training to determine the target model parameters; and update the target learning model based on the target model parameters.

[0130] Optionally, when the program is running, control the device where the non-volatile storage medium is located to perform the following steps: obtain an image of a target moving object to be recognized; input the image of the target moving object into a moving object re-identification model for analysis to obtain an image recognition result.

[0131] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.

[0132] In the above embodiments of the present application, the descriptions of the respective embodiments have their own focuses. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0133] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of units or modules can be in electrical or other forms.

[0134] The unit described as a separation component may or may not be physically separated, and the component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0135] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, may also be physically present individually for each unit, or two or more units may be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0136] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs, etc., which can store program codes.

[0137] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A model training method, comprising: obtaining a plurality of training data subsets corresponding to a plurality of perspectives, wherein each training data subset includes all moving object images under one perspective, and the moving object is an object that appears under the plurality of perspectives during movement; dividing the plurality of training data subsets into a first type of data set and a second type of data set; in each iteration step, inputting the first type of data set into a target learning model to train the target learning model based on a first loss function determined by initial model parameters, and determining an update amount of the initial model parameters; inputting the second type of data set into the target learning model, and training the target learning model with a second loss function determined based on the update amount of the initial model parameters to determine target model parameters; iteratively updating the target learning model according to the target model parameters determined in each iteration step wherein the determination process of the second loss function includes: respectively selecting a first training sample group from the first type of data set and a second training sample group from the second type of data set based on a preset rule; determining second model parameters based on the initial model parameters and the update amount of the initial model parameters; calculating a third loss function based on the moving object images in the second training sample group and the second model parameters, the third loss function having the same structure as the first loss function, and the parameters in the third loss function being the second model parameters; calculating a cross-perspective loss function based on the moving object images in the first training sample group, the moving object images in the second training sample group, and the second model parameters; and summing the third loss function and the cross-perspective loss function to obtain the second loss function.

2. The method according to claim 1, wherein, obtaining a plurality of training data subsets corresponding to a plurality of perspectives includes: obtaining a plurality of moving object images under a plurality of perspectives, and taking all the moving object images under each perspective as one of the training data subsets; wherein each of the moving object images has an identity label and a perspective label, the identity label is used to indicate the identity of the moving object corresponding to the moving object image, and the perspective label is used to indicate the perspective to which the moving object image belongs.

3. The method according to claim 1, wherein, dividing the plurality of training data subsets into a first type of data set and a second type of data set includes: randomly dividing the plurality of training data subsets into two parts during any training cycle of the training process to respectively obtain the first type of data set and the second type of data set; wherein the training process includes a plurality of the training cycles, and each training cycle includes a plurality of the iteration steps.

4. The method according to claim 1, wherein, inputting the first type of data set into a target learning model to train the target learning model based on a first loss function determined by initial model parameters and determining an update amount of the initial model parameters includes: selecting a first training sample group from the first type of data set based on a preset rule; Input the first training sample group into the target learning model, and determine the first loss function based on the initial model parameters of the target learning model; Determine the update amount of the initial model parameters based on the first loss function.

5. The method according to claim 4, wherein, Selecting the first training sample group from the first type of dataset based on a preset rule includes: Determining a first preset number of identity labels from the first type of dataset, and selecting a second preset number of moving object images from the multiple moving object images corresponding to each identity label to form the first training sample group.

6. The method according to claim 4, wherein, Inputting the first training sample group into the target learning model and determining the first loss function based on the initial model parameters of the target learning model includes: Calculating the first loss function based on the moving object images in the first training sample group and the initial model parameters of the target learning model, wherein the first loss function includes at least one of the following loss functions: cross-entropy loss function, contrastive loss function, triplet loss function, center loss function.

7. The method according to claim 5, wherein, Inputting the second type of dataset into the target learning model, and training the target learning model with a second loss function determined based on the update amount of the initial model parameters to determine the target model parameters, includes: Selecting a second training sample group from the second type of dataset based on a preset rule; Inputting the second training sample group into the target learning model, and determining the second loss function based on the update amount of the initial model parameters; Determining the target model parameters based on the second loss function.

8. The method according to claim 7, wherein, Selecting the second training sample group from the second type of dataset based on a preset rule includes: Determining a first preset number of identity labels from the second type of dataset, and selecting a second preset number of moving object images from the multiple moving object images corresponding to each identity label to form the second training sample group; wherein the first preset number of identity labels determined in the second type of dataset and the first type of dataset are exactly the same.

9. The method according to claim 1, wherein, Calculating a cross-view loss function based on the moving object images in the first training sample group, the moving object images in the second training sample group, and the second model parameters, includes: Analyzing the moving object features corresponding to all the moving object images in the first training sample group and the second training sample group, and calculating the mean of all the moving object features in the first training sample group; Calculating the similarity between each moving object feature in the second training sample group and the mean, and determining the cross-view loss function based on the similarity.

10. A method for re-identifying moving objects, including: Obtaining a target moving object image to be identified; Inputting the target moving object image into a moving object re-identification model for analysis to obtain an image recognition result; Among them, the moving object re-identification model is trained in the following way: Obtain multiple training data subsets corresponding to multiple perspectives, where each training data subset includes all moving object images under one perspective, and the moving object is an object that appears under the multiple perspectives during movement; Divide the multiple training data subsets into a first type of dataset and a second type of dataset; In each iteration step, input the first type of dataset into the target learning model, and train the target learning model based on the first loss function determined by the initial model parameters to determine the update amount of the initial model parameters; Input the second type of dataset into the target learning model, and train the target learning model with the second loss function determined based on the update amount of the initial model parameters to determine the target model parameters; Iteratively update the target learning model according to the target model parameters determined in each iteration step; Among them, the determination process of the second loss function includes: Select a first training sample group from the first type of dataset and a second training sample group from the second type of dataset respectively based on a preset rule; Determine the second model parameters based on the initial model parameters and the update amount of the initial model parameters; Calculate a third loss function based on the moving object images in the second training sample group and the second model parameters, where the third loss function has the same structure as the first loss function, and the parameters in the third loss function are the second model parameters; Calculate a cross-perspective loss function based on the moving object images in the first training sample group, the moving object images in the second training sample group, and the second model parameters; Sum the third loss function and the cross-perspective loss function to obtain the second loss function.

11. The method according to claim 10, wherein, Input the target moving object image into the moving object re-identification model for analysis to obtain an image recognition result, including: Determine the application mode of the current moving object re-identification model; The application mode includes a moving object authentication mode and a moving object retrieval mode. Among them, the moving object authentication mode is used to determine whether the moving objects in multiple target moving object images are the same moving object, and the moving object retrieval mode is used to retrieve all images of the moving object corresponding to the target moving object image from a database, where the database includes images of multiple moving objects under multiple perspectives; When it is determined that the current application mode is the moving object authentication mode, the moving object re-identification model analyzes the moving object features corresponding to multiple target moving object images, determines the similarity between multiple moving object features and compares it with a preset threshold, and determines whether the moving objects in multiple target moving object images are the same moving object according to the comparison result; When determining that the current application mode is the moving object retrieval mode, the moving object re-identification model analyzes the moving object features corresponding to the target moving object image, determines the similarity between the moving object features and the moving object features corresponding to all moving object images in the database, compares it with a preset threshold, and determines all images of the moving object corresponding to the target moving object image in the database according to the comparison result.

12. A model training device, comprising: an acquisition module, configured to acquire a plurality of training data subsets corresponding to a plurality of perspectives, wherein each training data subset includes all moving object images in one perspective, and the moving object is an object that appears in the plurality of perspectives during movement; a classification module, configured to divide the plurality of training data subsets into a first type of data set and a second type of data set; a determination module, configured to, in each iteration step, input the first type of data set into a target learning model to train the target learning model based on a first loss function determined by initial model parameters, and determine an update amount of the initial model parameters; input the second type of data set into the target learning model, and train the target learning model with a second loss function determined based on the update amount of the initial model parameters to determine target model parameters, wherein the determination process of the second loss function includes: respectively selecting a first training sample group from the first type of data set and a second training sample group from the second type of data set based on a preset rule; determining second model parameters based on the initial model parameters and the update amount of the initial model parameters; calculating a third loss function based on the moving object images in the second training sample group and the second model parameters, the third loss function having the same structure as the first loss function, and the parameters in the third loss function being the second model parameters; calculating a cross-perspective loss function based on the moving object images in the first training sample group, the moving object images in the second training sample group, and the second model parameters; summing the third loss function and the cross-perspective loss function to obtain the second loss function; an update module, configured to iteratively update the target learning model according to the target model parameters determined in each iteration step.

13. A non-volatile storage medium, wherein, the non-volatile storage medium includes a stored program, and when the program runs, it controls the device where the non-volatile storage medium is located to execute the model training method according to any one of claims 1 to 9.

14. An electronic device, comprising: a processor; and A memory, connected to the processor, for providing instructions for the processor to process the following processing steps: obtaining a plurality of training data subsets corresponding to a plurality of perspectives, where each training data subset includes all moving object images under one perspective, and the moving object is an object that appears under the plurality of perspectives during movement; dividing the plurality of training data subsets into a first type of data set and a second type of data set; in each iteration step, inputting the first type of data set into a target learning model to train the target learning model based on a first loss function determined by initial model parameters, and determining an update amount of the initial model parameters; inputting the second type of data set into the target learning model, and training the target learning model with a second loss function determined based on the update amount of the initial model parameters to determine target model parameters; iteratively updating the target learning model according to the target model parameters determined in each iteration step; Wherein, the determination process of the second loss function includes: respectively selecting a first training sample group from the first type of data set and a second training sample group from the second type of data set based on a preset rule; determining second model parameters based on the initial model parameters and the update amount of the initial model parameters; calculating a third loss function based on the moving object images in the second training sample group and the second model parameters, the third loss function having the same structure as the first loss function, and the parameters in the third loss function being the second model parameters; calculating a cross-perspective loss function based on the moving object images in the first training sample group, the moving object images in the second training sample group, and the second model parameters; summing the third loss function and the cross-perspective loss function to obtain the second loss function.

Citation Information

Patent Citations

  • Hybrid vehicle working condition prediction method based on meta-learning

    CN111047085A

  • Vehicle re-identification method in multi-view environment based on multi-center measurement loss

    CN111814584A