Model training method, population positioning method, device, equipment and medium

By selecting and training perspectives in multi-perspective crowd tasks and combining with pseudo-label perspectives, the problems of inaccurate prediction of scene-level crowds and high label costs in multi-perspective crowd tasks are solved, and higher prediction accuracy and lower label costs are achieved.

CN119540686BActive Publication Date: 2025-07-22SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510098294.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-07-22
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

There is a problem of inaccurate prediction of scene-level crowds in the existing multi-view crowd tasks, especially because the randomly sampled multi-view input causes large numbers of people in the scene to be missed, and the tag costs are high.

Method used

By selecting the first perspective and adding it to the labeled perspective set, crowd annotation is performed based on the labeled perspective set, the initial task model is trained, and the second perspective to the first perspective is gradually selected, and the pseudo-label perspective is used for training, and the target task model is constructed to reduce crowd omissions in the scene and reduce label costs.

Benefits of technology

It improves the prediction accuracy of multi-perspective tasks, reduces the omissions of crowds in the scene, reduces labeling costs, and enhances the generalization and performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119540686B_ABST
    Figure CN119540686B_ABST
Patent Text Reader

Abstract

The present application discloses a model training method, a crowd positioning method, a device, a device and a medium. The method includes selecting a first perspective based on a training data set and adding the first perspective to a set of labeled perspectives; training an initial task model based on a labeled training data set for annotating the first set of labeled perspectives; selecting a second perspective based on the initial task model and the set of labeled perspectives and adding the second perspective to the set of labeled perspectives, and retraining the initial task model until the nth perspective is selected; training a preset task model based on a labeled training data set for annotating the n perspectives. In the process of perspective selection, the present application fully considers the prediction results of the initial task model, so that the selected perspectives can cover more people in the scene, reduce the situation where people in the scene are missed, improve the detection accuracy of the target task model obtained by training, thereby improving the prediction accuracy of multi-perspective tasks and reducing the annotation cost of training the task model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and particularly to a model training method, a crowd positioning method, a device, a device and a medium. Background Art

[0002] Multi-view crowd tasks (such as multi-view crowd counting and multi-view crowd positioning, etc.) solve the problems of large-scale scene coverage, severe occlusion and ambiguity encountered in single-view tasks by fusing multiple calibrated camera views at the same time. However, existing multi-view crowd tasks mainly focus on task models and generally adopt randomly sampled multi-view inputs, and randomly sampled multi-views may have the problem of missing a large number of people in the scene, which will lead to inaccurate scene-level crowd predictions.

[0003] Therefore, the existing technology still needs to be improved. Summary of the Invention

[0004] The technical problem to be solved by this application is to provide a model training method, a crowd positioning method, a device, a device and a medium in view of the deficiencies of the existing technology.

[0005] To solve the above technical problem, the first aspect of this application provides a model training method, wherein the model training method specifically includes:

[0006] Obtain a training data set, where the training data set includes a number of multi-view images;

[0007] Select a first view based on the training data set, and add the first view to the labeled view set;

[0008] Perform crowd annotation on the training data set based on the labeled view set to obtain an annotated training data set, and train an initial task model corresponding to the crowd task based on the annotated training data set, where each labeled view in the annotated training data set carries a crowd label;

[0009] Select a second view based on the initial task model and the labeled view set, and add the second view to the labeled view set, and re-perform the step of performing crowd annotation on the training data set based on the labeled view set to obtain an annotated training data set until the view is selected, where is a positive integer;

[0010] Add the view to the labeled view set, perform crowd annotation on the training data set based on the labeled view set to obtain an annotated training data set, and train a preset task model corresponding to the crowd task based on the annotated training data set to obtain a target task model corresponding to the crowd task.

[0011] The described model training method, wherein, the process of selecting the perspective specifically includes:

[0012] Calculate the perspective evaluation value of each unlabeled perspective and the labeled perspective set on each frame of multi-perspective image in the training data set, wherein the perspective evaluation value is determined based on the scene coverage rate, the average perspective distance, and the perspective difference;

[0013] Calculate the sum of the perspective evaluation values of the perspective evaluation values corresponding to each unlabeled perspective, and select the perspective among the unlabeled perspectives based on the sum of the perspective evaluation values.

[0014] The described model training method, wherein the training of the initial task model corresponding to the crowd task based on the labeled training data set specifically includes:

[0015] For each frame of multi-perspective image in the training data set, select a pseudo-label perspective for the multi-perspective image;

[0016] Train the initial task model corresponding to the crowd task based on the single-perspective images corresponding to each labeled perspective and the pseudo-label perspective.

[0017] The described model training method, wherein the training of the initial task model corresponding to the crowd task based on the single-perspective images corresponding to each labeled perspective and the pseudo-label perspective specifically includes:

[0018] Input the single-perspective images corresponding to each labeled perspective and the pseudo-label perspective into the initial task model corresponding to the crowd task, and output the first predicted ground crowd density map through the initial task model;

[0019] Obtain the union of the fields of view of the labeled perspectives in the labeled perspective set to obtain the first combined field of view;

[0020] Determine the first ground crowd density map based on the first predicted ground crowd density map and the first combined field of view;

[0021] Construct the first loss function term based on the first ground crowd density map and the crowd labels of the labeled perspectives;

[0022] Train the initial task model based on the first loss function term.

[0023] The described model training method, wherein the training of the preset task model corresponding to the crowd task based on the labeled training data set to obtain the target task model corresponding to the crowd task specifically includes:

[0024] For each frame of multi-perspective image in the training data set, select in the From the unlabeled perspectives except the labeled perspective set, select pseudo-labeled perspectives, where , is a positive integer;

[0025] Based on the target perspectives and the single-perspective images corresponding to the pseudo-labeled perspectives, train a preset task model for the crowd task.

[0026] The model training method described above, where the training of the preset task model for the crowd task based on the target perspectives and the single-perspective images corresponding to the pseudo-labeled perspectives specifically includes:

[0027] Input the single-perspective images corresponding to the target perspectives and the pseudo-labeled perspectives into the preset task model, and output a second predicted ground crowd density map through the preset task model;

[0028] Obtain the union of the fields of view of each labeled perspective to obtain a first combined field of view, and the union of the fields of view of the target perspectives and the pseudo-labeled perspectives to obtain a second combined field of view;

[0029] Determine an intersecting field of view based on the first combined field of view and the second combined field of view;

[0030] Determine a second ground crowd density map based on the intersecting field of view and the crowd labels of each labeled perspective, and determine a third ground crowd density map based on the intersecting field of view and the second predicted ground crowd density map;

[0031] Construct a second loss function term based on the second ground crowd density map and the third ground crowd density map;

[0032] Train the preset task model based on the second loss function term.

[0033] The second aspect of the present application provides a crowd positioning method, using a target task model for the crowd task trained by using the model training method described above, where the crowd task is a crowd positioning task; the crowd positioning method specifically includes:

[0034] Obtain multi-perspective images, and input the multi-perspective images into the target task model;

[0035] Output a crowd positioning result through the target task model.

[0036] A third aspect of the present application provides a model training device, where the model training device specifically includes:

[0037] An acquisition module, configured to acquire a training data set, where the training data set includes a plurality of multi-view images;

[0038] A first selection module, configured to select a first view based on the training data set and add the first view to the labeled view set;

[0039] A first training module, configured to perform population annotation on the training data set based on the labeled view set to obtain an annotated training data set, and train an initial task model corresponding to the population task based on the annotated training data set, where each labeled view in the annotated training data set carries a population label;

[0040] A second selection module, configured to select a second view based on the initial task model and the labeled view set, add the second view to the labeled view set, and re-execute the step of performing population annotation on the training data set based on the labeled view set to obtain an annotated training data set until the view is selected, where is a positive integer;

[0041] A second training module, configured to add the view to the labeled view set, perform population annotation on the training data set based on the labeled view set to obtain an annotated training data set, and train a preset task model corresponding to the population task based on the annotated training data set to obtain a target task model corresponding to the population task.

[0042] A fourth aspect of the present application provides a computer-readable storage medium storing one or more programs, where the one or more programs can be executed by one or more processors to implement the steps in any of the above model training methods.

[0043] A fifth aspect of the present application provides a terminal device, which includes: a processor and a memory;

[0044] The memory stores a computer-readable program executable by the processor;

[0045] When the processor executes the computer-readable program, it implements the steps in any of the above model training methods.

[0046] Beneficial effects: Compared with the prior art, the present application provides a model training method, a crowd positioning method, a device, a device and a medium. The method includes selecting a first perspective based on the training data set and adding the first perspective to the labeled perspective set; performing crowd annotation on the training data set based on the labeled perspective set to obtain an annotated training data set, and training an initial task model corresponding to the crowd task based on the annotated training data set; selecting a second perspective based on the initial task model and the labeled perspective set, and adding the second perspective to the labeled perspective set, and re-executing the step of performing crowd annotation on the training data set based on the labeled perspective set to obtain an annotated training data set until the th perspective is selected; adding the th perspective to the labeled perspective set, performing crowd annotation on the training data set based on the labeled perspective set to obtain an annotated training data set, and training a preset task model corresponding to the crowd task based on the annotated training data set to obtain a target task model corresponding to the crowd task. In the process of perspective selection in the present application, the prediction results of the initial task model are fully considered, so that the selected perspective can cover more people in the scene, thereby reducing the situation where people in the scene are missed, and further improving the detection accuracy of the target task model obtained by training, and further improving the prediction accuracy of the multi-perspective task.

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0048] Figure 1 It is a flowchart of the model training method provided by the embodiment of the present application.

[0049] Figure 2 It is a principle flowchart of the model training method provided by the embodiment of the present application.

[0050] Figure 3 It is a principle flowchart of training an initial task model based on an annotated training data set.

[0051] Figure 4 It is a schematic diagram of scene coverage and ground crowd density.

[0052] Figure 5 It is a principle flowchart of training a preset task model based on an annotated training data set.

[0053] Figure 6 It is a principle block diagram of the model training device provided by the embodiment of the present application.

[0054] Figure 7 This is a block diagram of the terminal device provided by the embodiment of the present application. Detailed implementation manners

[0055] The embodiments of the present application provide a model training method, a crowd positioning method, an apparatus, a device, and a medium. To make the objectives, technical solutions, and effects of the present application clearer and more definite, the following further describes the present application in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0056] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0057] Those skilled in the art of the present technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless specifically defined as here.

[0058] It should be understood that the sequence numbers and magnitudes of the steps in this embodiment do not mean the order of execution. The order of execution of each process is determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0059] Through research, it is found that multi-view crowd tasks (such as multi-view crowd counting and multi-view crowd localization, etc.) solve the problems of large-scale scene coverage, severe occlusion, and ambiguity encountered in single-view tasks by fusing multiple calibrated camera views at the same time. However, existing multi-view crowd tasks mainly focus on the task model and generally adopt randomly sampled multi-view inputs. The randomly sampled multi-view method only focuses on the crowd prediction results of the input views, without paying attention to the scene-level crowd prediction results. This may lead to inaccurate scene-level crowd predictions due to a large number of scene-level people being omitted. Moreover, a large amount of labeled data is used each time the task model is trained, which requires a relatively high labeling cost.

[0060] To solve the above problems, in the embodiments of the present application, a first view is selected based on the training data set, and the first view is added to the labeled view set; the training data set is labeled for crowds based on the labeled view set to obtain a labeled training data set, and an initial task model corresponding to the crowd task is trained based on the labeled training data set; a second view is selected based on the initial task model and the labeled view set, and the second view is added to the labeled view set, and the step of labeling the training data set for crowds based on the labeled view set to obtain a labeled training data set is re-executed until the th view; the th view is added to the labeled view set, the training data set is labeled for crowds based on the labeled view set to obtain a labeled training data set, and a preset task model corresponding to the crowd task is trained based on the labeled training data set to obtain a target task model corresponding to the crowd task. In the process of view selection in the present application, the prediction results of the initial task model are fully considered, so that the selected views can cover more people in the scene, thereby reducing the situation of people in the scene being omitted, and further improving the detection accuracy of the trained target task model, and further improving the prediction accuracy of multi-view tasks.

[0061] At the same time, in the process of view selection in the present application and in the model training process after the th view is selected, pseudo-label views are introduced. By indirectly using the crowd labels of the labeled views as the crowd labels of the pseudo-label views, the task model can be trained using a large amount of unlabeled data under the condition of limited labels, improving the performance of the model and reducing the labeling cost.

[0062] The following further illustrates the application content through the description of the embodiments in conjunction with the accompanying drawings.

[0063] This embodiment provides a model training method, as shown in Figure 1 and Figure 2 The method includes:

[0064] S10. Obtain a training data set, where the training data set includes a number of multi-view images.

[0065] Specifically, the training data set includes a number of multi-view images. The number of multi-view images can be obtained by shooting the same scene or multiple scenes. Among them, the shooting times of the multi-view images corresponding to the same scene are different, and there may be multi-view images with the same shooting time among the multi-view images corresponding to different scenes. For example, multiple calibrated camera views can be deployed in one scene, and the scene can be shot at different times through the multiple calibrated camera views to obtain a number of multi-view images; or, multiple calibrated camera views can be deployed in multiple scenes, and then each scene can be shot separately to obtain a number of multi-view images.

[0066] It should be noted that when the training data set corresponds to multiple scenes, a certain number of views are selected for each scene, and then after all scenes have selected the same number of views, the target task model is trained based on the multi-view images corresponding to all scenes. The selection process of the same number of views for each scene and the processing process of each frame of multi-view image in the training process of the target task model are the same. Based on this, for the convenience of description, here, an example where the training data set corresponds to one scene is used to describe the subsequent selection process of the same number of views and the training process of the target task model. selection process of the same number of views and the training process of the target task model.

[0067] Furthermore, after the multi-view images are obtained by shooting, the data set composed of the multi-view images obtained by shooting can be directly used as the training data set, or the multi-view images obtained by shooting can be screened first, and then the data set composed of the screened multi-view images can be used as the training data set. In the embodiments of the present application, in order to improve the image quality of the multi-view images included in the training data set, after the multi-view images are obtained by shooting, the multi-view images obtained by shooting will be screened first, and then the data set composed of the screened multi-view images will be used as the training data set to further reduce the labeling cost. Among them, the process of screening the multi-view images obtained by shooting can be:

[0068] Obtain the largest view angle with the largest field of view in the scene corresponding to the captured image set;

[0069] Count the crowd counting results of the largest view angle in each frame of multi-view image in the captured image set, and use the multi-view image with the largest crowd counting result as the first frame of multi-view image;

[0070] Add the first frame of multi-view image to the selected image set, and remove the first frame of multi-view image from the captured image set;

[0071] Calculate the cosine similarity of each frame of multi-view images in the selected image set and each frame of multi-view images in the captured image set at the maximum viewing angle respectively, and select the second frame of multi-view images based on the calculated cosine similarity;

[0072] Add the second frame of multi-view images to the selected image set, and remove the second frame of multi-view images from the captured image set;

[0073] Calculate the cosine similarity of each frame of multi-view images in the selected image set and each frame of multi-view images in the captured image set at the maximum viewing angle respectively, and select the third frame of multi-view images based on the calculated cosine similarity;

[0074] And so on until the th frame of multi-view images is selected. Add the th frame of multi-view images to the selected image set, and use the selected image set as the training data set for the corresponding scene of the captured image set, where is a positive integer.

[0075] Specifically, the crowd counting result in the multi-view image can be obtained by using a preset single-view counting model (for example, the DM-Count model, etc.). That is to say, the crowd counting of the single-view image corresponding to the maximum viewing angle is counted by the single-view counting model, and the crowd counting is used as the crowd counting result in the multi-view image. After obtaining the crowd counting result corresponding to each frame of multi-view image, the multi-view image with the largest crowd counting result can be used as the first frame of multi-view image.

[0076] Furthermore, when selecting multi-view images based on the calculated cosine similarity, when there is only one frame of multi-view image in the selected image set, the multi-view image with the smallest calculated cosine similarity can be directly used as the second multi-view image, that is, the multi-view image with the smallest cosine similarity to the first frame of multi-view image in the captured image set at the maximum viewing angle is selected as the second multi-view image. When the selected image set includes multiple frames of multi-view images, after calculating all cosine similarities, for each frame of multi-view image in the captured image set, calculate the sum of its cosine similarities with each frame of multi-view image in the selected image set, and then use the multi-view image corresponding to the smallest sum of cosine similarities as the selected multi-view image. For example, when the selected image set includes the first frame of multi-view image and the second frame of multi-view image, the multi-view image with the smallest sum of the cosine similarity to the first frame of multi-view image and the cosine similarity to the second frame of multi-view image can be selected as the third multi-view image.

[0077] It should be noted that when taking multi - perspective images of multiple scenes, the captured dataset obtained for each scene can be screened for multi - perspective images according to the above process. Then, the set composed of the multi - perspective images screened for each scene is used as the training dataset.

[0078] This application selects frames of multi - perspective images from the captured dataset corresponding to the scene as training data. Then, only the frames of multi - perspective images in the scene need to be labeled. In this way, when the training requirements can be met, as few training data as possible are selected, and the training data that needs to be labeled can be calculated, thereby reducing the crowd - labeling cost.

[0079] S20. Select a first perspective based on the training dataset and add the first perspective to the set of labeled perspectives.

[0080] Specifically, the first perspective is one of the perspectives in the multi - perspective images. Among them, the first perspective is determined based on the crowd - counting results of each perspective in the multi - perspective images on the training dataset. The determination process can be as follows: First, obtain the crowd - counting results of each perspective in each frame of the multi - perspective images in the training dataset for each frame of the multi - perspective images. Then, sum up the crowd - counting results of each perspective on the selected F frames of multi - perspective images to obtain the total crowd - counting result corresponding to each perspective. Finally, select the perspective with the largest total crowd - counting result as the first perspective. Of course, in practical applications, other methods can also be used to select the first perspective, such as randomly selecting a perspective as the first perspective, or randomly selecting a perspective from the perspectives whose total crowd - counting results are greater than the preset crowd - counting quantity, etc.

[0081] In the embodiment of this application, by selecting the perspective with the largest total crowd - counting result as the first perspective, the first perspective can cover more people in the scene. And in the subsequent perspective selection, the initial task model trained based on the first perspective is used for selection, which can make the subsequently selected perspectives also cover more people in the scene, so that the selected perspectives can cover more people in the scene.

[0082] S30. Perform crowd annotation on the training dataset based on the set of labeled perspectives to obtain an annotated training dataset, and train an initial task model corresponding to the crowd task based on the annotated training dataset.

[0083] Specifically, the labeled view set includes the actively selected views. Performing population annotation on the training data set based on the labeled view set means performing population annotation on the single-view images corresponding to the labeled views in the labeled view set. For example, when the labeled view set includes the first view, the single-view image corresponding to the first view is annotated. Based on this, the multi-view images included in the annotated training data set are the same as the multi-view images included in the training data set. The single-view images corresponding to the labeled views in the annotated training data set carry population labels, while the single-view images of all views in the training data set do not carry population labels.

[0084] When training the initial task model corresponding to the population task based on the annotated training data set, the single-view images corresponding to the labeled views in the labeled view set can be directly used to train the initial task model corresponding to the population task. For example, the single-view images corresponding to the labeled views are input into the initial task model corresponding to the population task, and the predicted ground population density map is output through the initial task model corresponding to the population task. Then, a loss function term is constructed based on the predicted ground population density map and the population label corresponding to each single-view image, and the initial task model corresponding to the population task is trained based on the loss function term. Among them, the population task can be a population localization task, a population counting task, etc.

[0085] Furthermore, the labeled views with population labels and the views without population labels are in the same scene. The population labels of the labeled views can be used as pseudo-population labels for the views without population labels, so that the single-view images of the views without population labels can be added to the model training during the training process. Based on this, in one implementation, the training of the initial task model corresponding to the population task based on the annotated training data set specifically includes:

[0086] For each frame of multi-view image in the training data set, select a pseudo-label view for the multi-view image;

[0087] Train the initial task model corresponding to the population task based on the single-view images corresponding to each labeled view and pseudo-label view.

[0088] Specifically, the pseudo-label perspective is an unlabeled perspective among the multi-perspective images, that is, the pseudo-label perspective is included in the perspective set corresponding to the multi-perspective images, but not included in the labeled perspective set. Among them, the pseudo-label perspectives corresponding to each frame of multi-perspective images in the training dataset can be the same; or some of the pseudo-label perspectives corresponding to the multi-perspective images can be the same, and some can be different; or the pseudo-label perspectives corresponding to each frame of multi-perspective images can be all different. In the embodiments of the present application, an unlabeled perspective can be randomly selected from the unlabeled perspectives corresponding to each frame of multi-perspective images, and the randomly selected unlabeled perspective is used as the pseudo-label perspective corresponding to the multi-perspective image. Of course, in practical applications, other methods can also be used. For example, a frame of multi-perspective image is randomly selected from the training dataset, and then an unlabeled perspective is randomly selected from the unlabeled perspectives corresponding to the multi-perspective image, and the selected unlabeled perspective is used as the pseudo-label perspective corresponding to each frame of multi-perspective images, etc.

[0089] In the embodiments of the present application, by selecting a pseudo-label perspective for each frame of multi-perspective image, the initial task model for training the crowd task corresponding to the single-perspective images corresponding to the labeled perspectives in the labeled perspective set can be converted into the initial task model for training the crowd task corresponding to the single-perspective images corresponding to the labeled perspectives in the labeled perspective set and the single-perspective image corresponding to the pseudo-label perspective (for example, if the labeled perspective set includes perspectives, then perspectives will be used in model training and converted to using perspectives), which can improve the training speed of the initial task model. Of course, in practical applications, multiple pseudo-label perspectives can also be selected, such as 2, 3, etc.

[0090] Exemplarily, training the initial task model corresponding to the crowd task based on the augmented labeled training dataset specifically includes:

[0091] Input the single-perspective images corresponding to each labeled perspective and the pseudo-label perspective into the initial task model corresponding to the crowd task, and output the first predicted ground crowd density map through the initial task model;

[0092] Obtain the union of the fields of view of the labeled perspectives in the labeled perspective set to obtain the first combined field of view;

[0093] Determine the first ground crowd density map based on the first predicted ground crowd density map and the first combined field of view;

[0094] Construct the first loss function term based on the first ground crowd density map and the crowd labels of the labeled perspectives;

[0095] Train the initial task model based on the first loss function term.

[0096] Specifically, as Figure 3 shown, obtain the field of view of each labeled perspective, and then take the union of all fields of view to obtain the first combined field of view. After obtaining the first combined field of view, perform a masking operation on the first predicted ground crowd density map using the first combined field of view to obtain the first ground crowd density map. For example, multiply the first predicted ground crowd density map point-by-point with the first combined field of view to obtain the first ground crowd density map, such that the regions in the first predicted ground crowd density map that overlap with the first combined field of view remain unchanged, and the points in other regions become a preset value (such as 0, etc.).

[0097] After obtaining the first ground crowd density map, construct a first loss function term based on the first ground crowd density map and the crowd labels of the labeled perspectives, and then train the initial task model based on the first loss function term until the training process of the initial task model reaches a preset training metric. The preset training metric can be that the number of training times reaches a preset number, or that the model accuracy of the trained initial task model meets a preset requirement, etc. In addition, it should be noted that the construction method of the first loss function and the optimization process of the model parameters of the initial task model based on the first loss function can both adopt the process of optimizing the model parameters using the loss function in the training processes of existing crowd counting models and crowd localization models. At the same time, when training the initial task model, the multi-perspective images in the training dataset can also be divided into several training batches. Each training batch includes multi-perspective images with a certain amount of data. Each multi-perspective image determines the first ground crowd density map according to the above process, and then constructs the first loss function based on all the multi-perspective images in the training batch. For example, when the crowd task is a multi-perspective crowd counting task, the initial task model can use CVCS as the benchmark model and add FPN on the basis of a single-perspective decoder to capture the multi-scale information of the image; when the crowd task is a multi-perspective crowd localization task, the MVDet model can be used, etc.

[0098] S40. Select a second perspective based on the initial task model and the set of labeled perspectives, add the second perspective to the set of labeled perspectives, and re-execute the step of performing crowd annotation on the training dataset based on the set of labeled perspectives to obtain the annotated training dataset until the th perspective is selected.

[0099] Specifically, the initial task model for selecting the second perspective is the initial task model trained on the labeled training data set corresponding to the first perspective. That is, after training the initial task model based on the labeled training data set and when the initial task model meets the preset requirements, the second perspective is selected based on this initial task model and the set of labeled perspectives. Among them, the second perspective is included in all perspectives corresponding to the multi-perspective image, and it is different from the first perspective, that is, it is not included in the set of labeled perspectives.

[0100] After selecting the second perspective, the second perspective will be added to the set of labeled perspectives, and then the process of crowd annotation of the training data set based on the set of labeled perspectives will be re-executed to obtain the labeled training data set, so as to obtain the initial task model trained on the labeled training data set corresponding to the set of labeled perspectives with the second perspective added, and the third perspective is selected based on this initial task model and the set of labeled perspectives including the first perspective and the second perspective, and so on until the perspective, where is a positive integer. Among them, the determination of the labeled training data set and the training process of the initial task model after each perspective selection are the same as the above process, and will not be elaborated here.

[0101] Furthermore, represents the total number of actively selected perspectives, which can be preset, can be determined according to the training requirements of the target task model corresponding to the crowd task, or can be determined based on the scene crowd coverage rate (for example, when selecting the th perspective, obtain the ratio of the population covered by the single-perspective image of the th perspective to the population included in the scene. If the ratio reaches the preset ratio threshold, then is used as and the perspective selection is stopped, etc.). In addition, the selection process of each th perspective in 2, 3,..., is the same, and they are all selected based on the initial task model and the set of labeled perspectives. Here, the selection process of the

[0102] Exemplarily, the selection process of the th perspective specifically includes:

[0103] Calculate the perspective evaluation value of each unlabeled perspective and the set of labeled perspectives on each frame of multi-perspective image in the training data set;

[0104] Calculate the sum of the perspective evaluation values of the perspective evaluation values corresponding to each unlabeled perspective, and select the th perspective from the unlabeled perspectives based on the sum of the perspective evaluation values.

[0105] Specifically, the perspective evaluation value reflects the selection probability of the unlabeled perspective relative to the multi-perspective image, and the perspective evaluation value and the selection probability of the unlabeled perspective relative to the training data set. When selecting the perspective based on the perspective evaluation value among the unlabeled perspectives, the unlabeled perspective with the largest perspective evaluation value can be selected as the perspective. Among them, the perspective evaluation value is determined based on the scene coverage rate, the average perspective distance, and the perspective difference. The scene coverage rate is used to reflect the population coverage status of the perspective. The larger the scene coverage rate, the more people the perspective covers. Conversely, the smaller the scene coverage rate, the fewer people the perspective covers. The average perspective distance is used to reflect the distance relationship between the scene area and the image acquisition device (such as a camera). By the average perspective distance, the distance between the scene area and the image acquisition device meets the preset requirements to avoid the people in the multi-perspective image being too small, so as to ensure the clarity of the people in the multi-perspective image. The perspective difference is used to reflect the sparsity between perspectives, and the dispersion between perspectives is ensured through the perspective difference to avoid the selected perspectives clustering together.

[0106] In one implementation, the process of obtaining the perspective evaluation value may include:

[0107] For each unlabeled perspective, using the single-perspective image corresponding to the unlabeled perspective and the single-perspective images corresponding to each labeled perspective as input items, determining the ground population density map corresponding to the unlabeled perspective through the above initial task model, and determining the population mask map corresponding to the unlabeled perspective based on the ground population density map corresponding to the unlabeled perspective. Among them, the initial task model is trained after selecting the perspective, = 2, 3,..., ;

[0108] Determining the scene coverage rate corresponding to the unlabeled perspective based on the population mask map, determining the average perspective distance corresponding to the unlabeled perspective based on the ground population density map and the population mask map, and calculating the perspective difference corresponding to the unlabeled perspective;

[0109] Determining the perspective evaluation value of the unlabeled perspective based on the scene coverage rate, the average perspective distance, and the perspective difference.

[0110] Specifically, for the scene coverage rate, the higher the scene coverage rate, the more people the selected perspective can cover. Therefore, the scene coverage rate can be used as an evaluation item for perspective selection. The scene coverage rate can be determined based on the field of view of the perspective and the area of the scene. Specifically, let the th perspective in the union of the labeled perspective set and the unlabeled perspective have a field of view of , such as Figure 4As shown, the combined field of view of all labeled perspectives in the labeled perspective concentration is , and the scene has an area of . represents the width of the scene area, and represents the height of the scene area. Then the scene coverage rate can be expressed as . represents the sum of the number of labeled perspectives and the number of an unlabeled perspective in the labeled perspective concentration. The number of labeled perspectives in the labeled perspective concentration is .

[0111] In addition, in practical applications, due to possible occlusions in multi-perspective images, the field of view of a perspective for a scene cannot directly represent the true coverage range. Therefore, to improve the accuracy of the scene coverage rate. When determining the scene coverage rate, it can be based on the ground crowd density map predicted by the initial task model to determine the scene coverage rate. Specifically, first, the ground crowd density map predicted by the initial task model, which is used to reflect the probability distribution of people in the scene, , and then based on the ground crowd density map to determine the crowd mask map of the area where people are located . For example, as Figure 4 shown, applying a threshold (such as a pre-set threshold, etc.) on the ground crowd density map can obtain the crowd mask map . Finally, taking the crowd mask map as the combined field of view, and determining the scene coverage rate by calculating the ratio of the combined field of view to the area. Among them, the calculation formula of the scene coverage rate can be:

[0112]

[0113] Among them, represents the scene coverage rate.

[0114] For the average perspective distance, the average perspective distance can reflect the clarity of people in the single-perspective image corresponding to the perspective, and the average perspective distance can be used as an evaluation item for perspective selection. The average perspective distance can be the mean of the average distance values of the visible points of the labeled perspectives. Specifically, for each visible point on the crowd mask map , by taking the reciprocal sum of the perspective distances from this visible point in the world coordinate system to the perspectives that can see this visible point as the average distance value of the visible point , that is, the average distance value of the visible point , represents the average distance value of visible points , represents the viewing angle position in the world coordinate system, and when the visible point falls within the viewing field of the viewing angle , it participates in the calculation of the average distance value of the visible point . Therefore, the average viewing angle distance can be determined by calculating the mean value of all visible points in the combined viewing field of viewing angles, that is, the average viewing angle distance corresponding to the labeled viewing angle set can be , represents the average viewing angle distance corresponding to the labeled viewing angle set

[0115] In addition, since the ground crowd density map reflects the probability value of the appearance of people, the ground crowd density map can be used as the weight value of each point on the crowd mask map . The greater the probability, the more likely it is that people will appear at that point, and a higher weight is assigned to that point. Therefore, the average distance value of the visible point can also be expressed as .

[0116] Regarding the viewing angle difference, since multiple viewing angles can reduce the occlusion and ambiguity of objects to a great extent compared with a single viewing angle, the viewing angle difference reflects the sparsity between viewing angles, ensuring that the viewing angles are relatively dispersed and not clustered together. The viewing angle difference can be calculated by measuring the viewing angle similarity to obtain the viewing angle difference value. Among them, the expression of the viewing angle difference can be:

[0117]

[0118] where represents the viewing angle difference represents the viewing angle in the optical axis direction of the world coordinate system, that is, the direction vector represents the viewing angle in the optical axis direction of the world coordinate system, that is, the direction vector and are both constant terms used to avoid division by zero operations used to control the sensitivity of the viewing angle difference term

[0119] Furthermore, after obtaining the scene coverage rate, average viewing angle distance, and viewing angle difference, the product of the scene coverage rate, average viewing angle distance, and viewing angle difference can be used as the viewing angle evaluation value, that is, the viewing angle evaluation value can be expressed as 。

[0120] In the model training process of the embodiments of the present application, an active view selection training process will be carried out first. Specifically, after selecting the training data set and the first view corresponding to each training scenario, the first view is added to the labeled view set, and the initial task model is trained on the labeled view set and the training data set selected in each training scenario. When the training index of the initial task model reaches the set threshold, view selection starts. After a view is selected, the selected view is added to the labeled view set and the population label, and then the initial task model is retrained on the labeled view set and the training data set selected in each training scenario, and so on until the th view ( the number of pre-labeled views). In the active view selection training process of the present application, the view selection process and the model training process are combined, so that the task model can guide the view selection process, which can improve the accuracy of view selection, and at the same time, the selected views can ensure the accuracy of the prediction results of the task model. Moreover, in the process of guiding view selection through the task model, the view evaluation value determined based on the scene coverage rate, the average view distance, and the view difference is used as the basis for view selection, fully considering the geometric information of the scene and the view, the crowd position information, and the crowd probability distribution information, and can select views that cover more people in the scene, so that the labeled views can be better used for model training and model testing, and better scene-level crowd prediction results can be obtained.

[0121] S50. Add the th view to the labeled view set, perform population annotation on the training data set based on the labeled view set to obtain an annotated training data set, and train a preset task model corresponding to the population task based on the annotated training data set to obtain a target task model corresponding to the population task.

[0122] Specifically, after the th view, the view selection process is completed, and only the training process needs to be carried out subsequently. For this reason, after adding the th view to the labeled view set, label the population label for the th view, and use the training data set of the th view as the final annotated training data set, and train a preset task model corresponding to the population task based on this annotated training data set. Among them, the preset task model can be the initial task model trained using the annotated training data set of the th view, or it can be a preset default task model, etc. In the embodiments of the present application, the The initial task model trained with the labeled training dataset for perspective annotation is used as the preset task model, and then the preset task model is trained based on the labeled training dataset to obtain the target task model.

[0123] Furthermore, when training the preset task model corresponding to the crowd task based on the labeled training dataset, the single-perspective images corresponding to perspectives can be directly used as the input items of the preset task model. Then, the predicted ground crowd density map is determined through the preset task model, and then a loss function item is constructed based on the predicted ground crowd density map and the crowd label. The model parameters of the preset task model are optimized through the loss function item to obtain the target task model.

[0124] However, in practical applications, since perspectives are only part of the multi-perspective images, and there are other unlabeled perspectives in the multi-perspective images. Therefore, during the training process after selecting perspectives, the single-perspective images corresponding to the unlabeled perspectives can be included in the training process of the preset task model, and training labels are generated for the input items including the unlabeled perspectives to improve the model performance and generalization ability of the target training model. Specifically, considering that the labeled perspectives and the unlabeled perspectives are in the same scene, the crowd labels of the labeled perspectives can be used to construct pseudo-crowd labels for the unlabeled perspectives.

[0125] Based on this, the process of training the preset task model corresponding to the crowd task based on the labeled training dataset to obtain the target task model corresponding to the crowd task specifically includes:

[0126] For each frame of multi-perspective image in the training dataset, target perspectives are selected from the labeled perspective set, and pseudo-label perspectives are selected from the unlabeled perspectives except the labeled perspective set, where ,, is a positive integer;

[0127] Based on the single-perspective images corresponding to the target perspectives and pseudo-label perspectives, the preset task model corresponding to the crowd task is trained.

[0128] Specifically, the target perspectives can be the labeled perspectives in the labeled perspective set. For example, labeled perspectives can be randomly selected from the labeled perspective set as the target perspectives. After selecting the target perspectives, pseudo-label perspectives are selected from the unlabeled perspectives, so that the number of single-perspective images used to train the preset task model is . Then, single-view images corresponding to the target viewpoints and pseudo-label viewpoints are used as input items of the preset task model, and the preset task model is trained through single-view images corresponding to the target viewpoints and pseudo-label viewpoints.

[0129] In one implementation, as Figure 5 shown, training the preset task model corresponding to the crowd task based on the augmented labeled training dataset specifically includes:

[0130] Input single-view images corresponding to the target viewpoints and pseudo-label viewpoints into the preset task model, and output a second predicted ground crowd density map through the preset task model;

[0131] Obtain the union of the fields of view of each labeled viewpoint to obtain a first combined field of view, and the union of the fields of view of the target viewpoints and pseudo-label viewpoints to obtain a second combined field of view;

[0132] Determine an intersecting field of view based on the first combined field of view and the second combined field of view;

[0133] Determine a second ground crowd density map based on the intersecting field of view and the crowd labels of each labeled viewpoint, and determine a third ground crowd density map based on the intersecting field of view and the second predicted ground crowd density map;

[0134] Construct a second loss function term based on the second ground crowd density map and the third ground crowd density map;

[0135] Train the preset task model based on the second loss function term.

[0136] Specifically, the determination processes of the first combined field of view and the second combined field of view are the same as the determination process of the first combined field of view above, and will not be elaborated here. The intersecting field of view is the intersection area of the first combined field of view and the second combined field of view. That is to say, the overlapping area of the first combined field of view and the second combined field of view can be obtained, and the obtained overlapping area is used as the intersecting field of view. In addition, the second ground crowd density map, the third ground crowd density map, the second loss function term, and training the preset task model based on each second loss function term are the same as those in the above viewpoint selection process, and will not be elaborated here.

[0137] In the embodiment of the present application, unlabeled perspectives are added both in the active perspective selection training process and the individual model training process, and the labeled perspectives are used to construct pseudo-population labels for the input items with unlabeled perspectives added, realizing the use of a small number of labeled perspectives to provide labels for a large number of unlabeled perspectives, enabling a large number of unlabeled perspectives to be used for model training and reducing the labeling cost.

[0138] In addition, after the target task model is trained, the target task model will also be tested. Specifically, as Figure 2 shown, the testing process includes an active perspective selection testing process and test label generation. Among them, the active perspective selection testing process is the same as the above-mentioned active perspective selection training process. In test label generation, the test metrics consider all the people present in the scene (i.e., scene-level labels), rather than only the people appearing in the selected perspectives as before. This can not only reflect the quality of the selected perspectives but also the quality of model training, and all the frames of the test scene are used for perspective selection during the perspective selection process.

[0139] In summary, the present embodiment provides a model training method. The method includes selecting a first perspective based on the training data set and adding the first perspective to the labeled perspective set; performing population annotation on the training data set based on the labeled perspective set to obtain an annotated training data set, and training an initial task model corresponding to the population task based on the annotated training data set; selecting a second perspective based on the initial task model and the labeled perspective set, and adding the second perspective to the labeled perspective set, and re-performing the step of performing population annotation on the training data set based on the labeled perspective set to obtain an annotated training data set until the perspective is selected; adding the perspective to the labeled perspective set, performing population annotation on the training data set based on the labeled perspective set to obtain an annotated training data set, and training a preset task model corresponding to the population task based on the annotated training data set to obtain a target task model corresponding to the population task. The present application fully considers the prediction results of the initial task model during the perspective selection process, enabling the selected perspectives to cover more people in the scene, thereby reducing the situation where people in the scene are missed, and further improving the detection accuracy of the target task model obtained by training, and further improving the prediction accuracy of the multi-perspective task.

[0140] To further illustrate the better generalization and scene-level crowd prediction results obtained by the model training method provided in the embodiments of the present application through the combination of active view selection and model training. In the embodiments of the present application, crowd tasks are taken as examples of multi-view crowd counting tasks and multi-view crowd localization tasks for comparative experiments. In the multi-view crowd counting task, the results of several methods under the same baseline model with different training methods, different view selection methods, and the addition of pseudo-label views are compared. As shown in Table 1, it is proved that the method provided in the embodiments of the present application can select better views for model training and testing, obtaining better scene-level crowd prediction results. It is also proved that the method provided in the embodiments of the present application can utilize a large amount of unlabeled data for training when only using a small amount of labeled data, enhancing the generalization and performance of the model. In the multi-view crowd localization task, the results of using different view selection methods and pseudo-label views under the same model are compared. As shown in Tables 2 and 3, where the multi-view selection method (MVSelect) uses all labeled data, it is proved that the method provided in the embodiments of the present application can achieve good results when only using a small amount of labeled data, reducing the labeling cost. It is also proved that the method provided in the embodiments of the present application considers the effectiveness of the active view selection framework of scene and view geometric information, crowd position information, and crowd probability distribution information.

[0141] Table 1 Schematic diagram of the results of the CVCS dataset for the multi-view crowd counting task

[0142]

[0143] Table 2 Schematic diagram of the results of the MultiviewX dataset for the multi-view crowd localization task

[0144]

[0145] Table 3 Schematic diagram of the results of the Wildtrack dataset for the multi-view crowd localization task

[0146]

[0147] Based on the above model training method, this embodiment provides a crowd localization method, which uses a target task model corresponding to a crowd task trained by using the above-mentioned model training method, and the crowd task is a crowd localization task; the crowd localization method specifically includes:

[0148] Obtain multi-view images, and input the multi-view images into the target task model;

[0149] Output the crowd localization result through the target task model.

[0150] Based on the above model training method, this embodiment provides a model training device, asFigure 6 As shown in the figure, the model training device specifically includes:

[0151] An acquisition module 100, configured to acquire a training data set, where the training data set includes a plurality of multi-view images;

[0152] A first selection module 200, configured to select a first view based on the training data set and add the first view to the labeled view set;

[0153] A first training module 300, configured to perform population annotation on the training data set based on the labeled view set to obtain an annotated training data set, and train an initial task model corresponding to the population task based on the annotated training data set, where each labeled view in the annotated training data set carries a population label;

[0154] A second selection module 400, configured to select a second view based on the initial task model and the labeled view set, add the second view to the labeled view set, and re-execute the step of performing population annotation on the training data set based on the labeled view set to obtain an annotated training data set until the th view is selected, where is a positive integer;

[0155] A second training module 500, configured to add the th view to the labeled view set, perform population annotation on the training data set based on the labeled view set to obtain an annotated training data set, and train a preset task model corresponding to the population task based on the annotated training data set to obtain a target task model corresponding to the population task.

[0156] Based on the above model training method, this embodiment provides a computer-readable storage medium, where the computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the model training method as described in the above embodiment.

[0157] Based on the above model training method, this application further provides a terminal device, as Figure 7 shown, which includes at least one processor 20; a display screen 21; and a memory 22, and may further include a communication interface 23 and a bus 24. Among them, the processor 20, the display screen 21, the memory 22, and the communication interface 23 can complete mutual communication through the bus 24. The display screen 21 is set to display a user guidance interface preset in the initial setting mode. The communication interface 23 can transmit information. The processor 20 can call the logical instructions in the memory 22 to execute the method in the above embodiment.

[0158] In addition, when the logic instructions in the above-mentioned memory 22 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.

[0159] As a computer-readable storage medium, the memory 22 can be set to store software programs and computer-executable programs, such as the program instructions or modules corresponding to the methods in the embodiments of the present disclosure. The processor 20 executes functional applications and data processing by running the software programs, instructions or modules stored in the memory 22, that is, the methods in the above embodiments are implemented.

[0160] The memory 22 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal device, etc. In addition, the memory 22 may include a high-speed random access memory and may also include a non-volatile memory. For example, various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs, can also be transient storage media.

[0161] In addition, the specific processes of loading and executing multiple instruction processors in the above storage medium and terminal device have been described in detail in the above methods, and will not be elaborated here one by one.

[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A model training method, characterized in that, The described model training method specifically includes: Obtain a training data set, where the training data set includes a number of multi-view images; Select a first view based on the training data set and add the first view to the set of labeled views; Perform crowd annotation on the training data set based on the set of labeled views to obtain an annotated training data set, and train an initial task model corresponding to the crowd task based on the annotated training data set, where each labeled view in the annotated training data set carries a crowd label; Select a second perspective based on the initial task model and the set of labeled perspectives, add the second perspective to the set of labeled perspectives, and re - execute the step of performing crowd annotation on the training data set based on the set of labeled perspectives to obtain an annotated training data set until the th perspective is selected, where is a positive integer; Add the viewpoint to the labeled viewpoint set, perform crowd annotation on the training data set based on the labeled viewpoint set to obtain an annotated training data set, and train a preset task model corresponding to the crowd task based on the annotated training data set to obtain a target task model corresponding to the crowd task; Among them, the process of selecting the viewing angle specifically includes: For each unlabeled perspective, using the single-perspective image corresponding to the unlabeled perspective and the single-perspective images corresponding to each labeled perspective as input items, determine the ground crowd density map corresponding to the unlabeled perspective through the above initial task model, and determine the crowd mask map corresponding to the unlabeled perspective based on the ground crowd density map corresponding to the unlabeled perspective, where the initial task model is trained after selecting the perspective, = 2, 3,..., ; Determine the scene coverage rate corresponding to the unlabeled view based on the crowd mask map, determine the average view distance corresponding to the unlabeled view based on the ground crowd density map and the crowd mask map, and calculate the view difference corresponding to the unlabeled view; Determine the view evaluation value of the unlabeled view based on the scene coverage rate, average view distance, and view difference; Calculate the sum of the perspective evaluation values corresponding to each unlabeled perspective, and select the perspective based on the sum of the perspective evaluation values among the unlabeled perspectives.

2. The model training method according to claim 1, wherein The training of the initial task model corresponding to the crowd task based on the annotated training data set specifically includes: For each multi-view image in the training data set, select a pseudo-label view for the multi-view image; Train an initial task model corresponding to the crowd task based on the single-view images corresponding to each labeled view and the pseudo-label view.

3. The model training method according to claim 2, wherein The training of the initial task model corresponding to the crowd task based on the single-view images corresponding to each labeled view and the pseudo-label view specifically includes: Input the single-view images corresponding to each labeled view and the pseudo-label view into the initial task model corresponding to the crowd task, and output a first predicted ground crowd density map through the initial task model; Obtain the union of the fields of view of the labeled views in the set of labeled views to obtain a first combined field of view; Determine a first ground crowd density map based on the first predicted ground crowd density map and the first combined field of view; Construct a first loss function term based on the first ground crowd density map and the crowd labels of the labeled views; Train the initial task model based on the first loss function term.

4. The model training method according to claim 1, wherein The training of the preset task model corresponding to the crowd task based on the annotated training data set to obtain the target task model corresponding to the crowd task specifically includes: For each frame of multi-view image in the training dataset, select target views from the labeled view set, and select pseudo-label views from the unlabeled views other than the labeled view set, where , is a positive integer; Based on the above-mentioned target perspectives and the single-perspective images corresponding to the pseudo-label perspectives, train a preset task model for the crowd task.

5. The model training method according to claim 4, wherein The one based on the target perspectives and the preset task model corresponding to the single-perspective image training crowd task corresponding to the pseudo-label perspectives specifically includes: Input the single-view images corresponding to target viewpoints and pseudo-label viewpoints into a preset task model, and output a second predicted ground crowd density map through the preset task model; Obtain the union of the visual fields of each labeled perspective to obtain a first combined visual field, and the target perspective and union of the visual fields of the pseudo-labeled perspectives to obtain a second combined visual field; Determine an intersecting field of view based on the first combined field of view and the second combined field of view; Determine a second ground crowd density map based on the intersecting field of view and the crowd labels of each labeled view, and determine a third ground crowd density map based on the intersecting field of view and the second predicted ground crowd density map; Construct a second loss function term based on the second ground crowd density map and the third ground crowd density map; Train the preset task model based on the second loss function term.

6. A method for crowd positioning, characterized in that, Use the target task model corresponding to the crowd task trained by using the model training method described in any one of claims 1-5, where the crowd task is a crowd localization task; the crowd localization method specifically includes: Obtain a multi-view image and input the multi-view image into the target task model; Output a crowd localization result through the target task model.

7. A model training device, characterized in that, The described model training device specifically includes: An acquisition module for acquiring a training data set, where the training data set includes a number of multi-view images; A first selection module, configured to select a first perspective based on the training data set and add the first perspective to the labeled perspective set; A first training module, configured to perform crowd annotation on the training data set based on the labeled perspective set to obtain an annotated training data set, and train an initial task model corresponding to the crowd task based on the annotated training data set, wherein each labeled perspective in the annotated training data set carries a crowd label; A second selection module, configured to select a second perspective based on the initial task model and the set of labeled perspectives, add the second perspective to the set of labeled perspectives, and re-execute the step of performing crowd annotation on the training data set based on the set of labeled perspectives to obtain an annotated training data set until the th perspective is selected, where The second training module is used to add the perspective to the labeled perspective set, perform population annotation on the training data set based on the labeled perspective set to obtain an annotated training data set, and train a preset task model corresponding to the population task based on the annotated training data set to obtain a target task model corresponding to the population task; Among them, the process of selecting the viewing angle specifically includes: For each unlabeled perspective, using the single-perspective image corresponding to the unlabeled perspective and the single-perspective images corresponding to each labeled perspective as input items, the ground crowd density map corresponding to the unlabeled perspective is determined through the above initial task model, and the crowd mask map corresponding to the unlabeled perspective is determined based on the ground crowd density map corresponding to the unlabeled perspective, where the initial task model is trained after selecting the perspective, = 2, 3,..., ; Determine the scene coverage rate corresponding to the unlabeled perspective based on the crowd mask map, determine the average perspective distance corresponding to the unlabeled perspective based on the ground crowd density map and the crowd mask map, and calculate the perspective difference corresponding to the unlabeled perspective; Determine the perspective evaluation value of the unlabeled perspective based on the scene coverage rate, the average perspective distance, and the perspective difference; Calculate the sum of the perspective evaluation values corresponding to each unlabeled perspective, and select the perspective based on the sum of the perspective evaluation values among the unlabeled perspectives.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the model training method according to any one of claims 1-5.

9. A terminal device, characterized in that, Including: A processor and a memory; The memory stores a computer-readable program executable by the processor; When the processor executes the computer-readable program, the steps in the model training method according to any one of claims 1-5 are implemented.