Meta-learning based cross-domain small sample crowd counting method

By employing a meta-learning-based cross-domain small sample population counting method, the source model is trained using a synthetic dataset, and a scene adaptation module and knowledge distillation regularization fine-tuning are introduced to solve the problems of data dependence and cross-domain adaptation, achieving low-cost and high-efficiency population counting.

CN118470630BActive Publication Date: 2026-01-23FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410535056.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-30
Publication Date
2026-01-23
Estimated Expiration
2044-04-30

AI Technical Summary

Technical Problem

Existing population counting methods rely on large amounts of labeled data, resulting in high costs for data collection and labeling, and insufficient generalization ability in cross-domain scenarios.

Method used

We employ a meta-learning-based cross-domain small sample population counting method, train the source model using a synthetic dataset, and reduce data requirements and improve model generalization ability through a scenario adaptation module and knowledge distillation regularization fine-tuning technique.

Benefits of technology

It significantly reduces the cost of data collection and annotation, while improving the model's adaptability and accuracy in cross-domain scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118470630B_ABST
    Figure CN118470630B_ABST
Patent Text Reader

Abstract

The application provides a meta-learning-based cross-domain small sample crowd counting method, which comprises the following steps: step S1: sampling and constructing a meta-training set and a meta-test set for a synthetic crowd counting dataset and a real-world crowd counting dataset respectively; step S2: constructing a scene adaptive module for dynamically adjusting a network according to scene context information; step S3: constructing a crowd counting model; step S4: designing a meta-learning-based training strategy; and step S5: testing the model obtained in step S4 on the meta-test set by using a knowledge distillation-based regularization fine-tuning method. The application significantly reduces the cost of data collection and labeling, solves the overfitting problem by using the knowledge distillation-based regularization fine-tuning method, and further introduces the scene adaptive module for adjusting the network according to the scene context information into the model to improve the generalization ability of the model in the cross-domain situation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing and computer vision, and particularly to a cross-domain small sample crowd counting method based on meta learning. BACKGROUND

[0002] With the continuous advancement of urbanization and the increase in the number of urban residents, public places such as airports, parks and shopping malls often gather a large number of people. However, the risk of trampling caused by crowd congestion brings great challenges to social safety governance.

[0003] In order to solve the problems caused by crowd congestion, the safety management personnel of public places need to monitor and analyze the crowd dynamics in real time, and once the crowd density exceeds the reasonable range, measures should be taken immediately to evacuate to ensure public safety. However, this work requires a large amount of human resources, thereby greatly increasing the management cost. Therefore, in order to improve work efficiency, more and more methods in the field of crowd analysis begin to use computer vision technology instead of manual monitoring. Crowd counting, as one of the basic tasks of crowd analysis, aims to obtain the total number and density distribution of people in a specific area in an image or video, providing data support for subsequent crowd analysis and management.

[0004] At present, most crowd counting methods are based on the fully supervised learning paradigm under traditional deep learning, that is, they rely on a large number of labeled data sets for training, and require the distribution of training and test data to be consistent, otherwise the performance of these methods may be greatly reduced. However, the cost of collecting and labeling a large amount of crowd data is very high. Therefore, how to reduce the dependence of crowd counting algorithms on data collection and labeling has become an important problem in the research of crowd counting. In recent years, to cope with the challenge of data cost, methods such as unsupervised domain adaptation crowd counting, small sample crowd counting and domain generalization crowd counting have emerged. However, these methods still have limitations: the former two require a large amount of data with the same distribution as the target scene for training, which is costly. The latter requires training on a variety of real data, which also involves a large amount of data collection and labeling. Therefore, the data dependence problem of crowd counting is worth further research. SUMMARY

[0005] The present application proposes a cross-domain small sample crowd counting method based on meta learning, which uses synthetic crowd counting data sets to train a source model, and uses a small amount of sample data of the target scene to train the source model, so that the model is applicable to the scene, significantly reducing the cost of data collection and labeling. And based on meta learning, an algorithm framework is proposed, including pre-training, meta training and meta testing stages, among which the meta testing stage uses a regularization fine-tuning method based on knowledge distillation to solve the problem of overfitting. In addition, a scene adaptive module is introduced in the model to adjust the network according to the scene context information, which is used to improve the generalization ability of the model in the cross-domain situation.

[0006] The application adopts the following technical solutions.

[0007] The meta-learning-based cross-domain small-sample crowd counting method comprises the following steps.

[0008] Step S1: sample and construct a meta-training set and a meta-test set for the synthetic crowd counting dataset and the real-world crowd counting dataset respectively;

[0009] Step S2: construct a scene adaptive module for dynamically adjusting the network according to the scene context information;

[0010] Step S3: construct a crowd counting model;

[0011] Step S4: design a meta-learning-based training strategy;

[0012] Step S5: use the model obtained in step S4 to test on the meta-test set by using a knowledge distillation-based regularization fine-tuning method.

[0013] Step S1 specifically comprises the following steps.

[0014] Step S11: first divide the synthetic crowd counting dataset according to the scene, randomly select a scene, and randomly extract a plurality of data in the scene as a support set and a query set to form a small-sample task, wherein the support set is used for model learning, and the query set is used for evaluating the learning effect of the model;

[0015] Repeat this step to form a meta-training set D train ; the meta-training set is formed in units of the small-sample task P in step S11, that is, wherein T train is the number of small-sample tasks in the meta-training set; and wherein is the support set of the i-th small-sample task in the meta-training set, is the query set of the i-th small-sample task in the meta-training set;

[0016] Step S12: first randomly select a scene from the real-world crowd counting dataset, sample data in the scene to form a support set and a query set, and form a small-sample task, wherein the support set is used for model learning, and the query set is used for testing the effect of the model; repeat this step to form a meta-test set D test ; the meta-test set is formed in units of the small-sample task P in step S12, that is, wherein T test is the number of small-sample tasks in the meta-test set; and wherein is the support set of the i-th small-sample task in the meta-test set, is the query set of the i-th small sample task in the meta-test set.

[0017] Step S2 specifically comprises the following steps:

[0018] Step S21: input the feature F in from the pre-convolutional layer of the scene adaptive module into the scene adapter, and the specific expression is:

[0019] λ = softmax(w2(w1(Avg(F in ))+b1)+b2),

[0020] wherein λ represents the coefficient of the convolution kernel template obtained by the scene adapter, and the dimension size is 4, Avg(·) represents the average pooling layer, w i and b i are the weights and bias of the i-th full connection layer, and softmax(·) represents the softmax layer; step S22: multiply the coefficient obtained in step S21 by the convolution kernel template to obtain the scene adaptive convolution, and the specific expression is:

[0021]

[0022] wherein Θ d represents the scene adaptive convolution, V n represents the n-th convolution kernel template, and λ n represents the coefficient corresponding to the n-th convolution kernel template. The dimension of the convolution kernel template is C in ×1×k×k, C in ×C out ×1×1, C in ×C out ×k×1 and C in ×C out ×1×k; wherein C in , C out and k respectively represent the input channel number, the output channel number and the convolution kernel size of the original convolutional layer at the embedded position of the scene adaptive module;

[0023] Step S23: use the scene adaptive convolution obtained in step S22 to extract scene-specific features, which are used to enhance the features extracted by the original convolutional layer at the embedded position, to obtain the enhanced features F out , and the specific expression is:

[0024] F out = Θ(F in )+ Θ d (F in ),

[0025] where Θ is the original convolutional layer.

[0026] In step S3, the crowd counting model comprises a feature extraction network, a density map regression network and a scene adaptive module.

[0027] Step S3 specifically comprises the following steps:

[0028] Step S31: first use the first 10 layers of VGG-16 as the feature extraction network to extract high-level features of the current image;

[0029] Step S32: input the high-level features obtained in step S31 into the density map regression network, which is composed of convolutional layers, and embed the scene adaptive module in a residual manner on each convolutional layer.

[0030] In step S4, the training strategy is divided into two stages: pre-training and meta-training.

[0031] Step S4 specifically comprises the following steps:

[0032] Step S41: train the crowd counting model constructed in step S3 on the synthetic crowd counting dataset in batches with mean square error as the loss function until the loss value converges and tends to be stable, and obtain an initial model;

[0033] Step S42: train the initial model obtained in step S41 on the meta-training set D train constructed in step S11; first fix the parameters of the feature extraction network of the initial model obtained in step S41, and train the model on D train The inner loop is executed first, that is, the support set is used to obtain the task-specific parameters The specific expression is:

[0034]

[0035] where L is the mean square error loss, and the formula is:

[0036] L(x; θ) = ||f θ (x) - y|| 2 ,

[0037] where f θ (·) refers to the crowd counting model with parameters θ, x is the input crowd image, and y is the density map label corresponding to the crowd image;

[0038] Step S43: use the task-specific parameters obtained in step S42 to execute the outer loop; the query set The loss is calculated using the learning rate β to update the model parameters θ. The specific expression is as follows:

[0039]

[0040] Step S44: Repeat steps S42 to S43 on a small sample task basis until all small sample tasks in the meta-training set have been traversed.

[0041] In step S42, the task-specific parameters are obtained from the model parameters θ through 10 gradient descent iterations.

[0042] Step S5 specifically includes the following steps:

[0043] Step S51: In the test meta test set D test The i-th task First, the student model and teacher model are created by copying the model obtained in step S4; the student model is updated by the supervised branch and the distillation branch; the loss function consists of the supervised loss and the distillation loss, and the specific expression is as follows:

[0044] L T =L S +ρL D ,

[0045] Where ρ is the weight of the distillation loss, L S It is a monitoring loss, L D It is the distillation loss; on the supervised branch, the support set is used. Density map labels Calculate the supervised loss L on the support set S The specific expression is:

[0046]

[0047] in The representative parameter is θ S The student model; on the distillation branch, the support set will be... After the image is input into the teacher model, a density map result is generated and used as a pseudo-label. Next, the pseudo-label is used to supervise the student model, and the distillation loss L is obtained. D The specific expression is:

[0048]

[0049] Step S52: Update the teacher model by applying an exponential moving average to the parameters of the student model. The specific expression is as follows:

[0050] θ t =ηθt +(1-η)θ S ,

[0051] Where η is the attenuation factor;

[0052] Step S53: Repeat steps S51 to S52 multiple times to obtain the task. Specific parameters are used in the query set Test on;

[0053] Step S54: Repeat steps S51 to S53 until all small sample tasks on the meta-test set have been traversed, average the test results of all small sample tasks, and obtain the final evaluation result of the model.

[0054] In step S53, the task is obtained after repeating steps S51 to S52 twenty times. Specific parameters.

[0055] This invention trains a source model using a synthetic crowd counting dataset, and then trains the source model using a small amount of sample data from the target scene, making the model suitable for that scene and significantly reducing the cost of data collection and annotation. An algorithmic framework based on meta-learning is proposed, including pre-training, meta-training, and meta-testing stages. The meta-testing stage uses a knowledge distillation-based regularization fine-tuning method to solve the overfitting problem. Furthermore, a scene adaptation module is introduced into the model to adjust the network based on scene context information, improving the model's generalization ability in cross-domain scenarios. Attached Figure Description

[0056] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:

[0057] Appendix Figure 1 This is a schematic diagram of the overall process of the method according to an embodiment of the present invention;

[0058] Appendix Figure 2 This is a schematic diagram of the network model structure in an embodiment of the present invention;

[0059] Appendix Figure 3 This is a schematic diagram of the cross-domain small sample population counting method based on meta-learning in an embodiment of the present invention;

[0060] Appendix Figure 4 This is a schematic diagram of the regularization fine-tuning method based on knowledge distillation in an embodiment of the present invention. Detailed Implementation

[0061] To make the features and advantages of this patent more apparent and understandable, specific embodiments are provided below for detailed explanation:

[0062] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0063] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0064] The cross-domain small sample population counting method based on meta-learning provided in this embodiment, such as Figure 1 , Figure 2 , Figure 3 , Figure 4 As shown,

[0065] A meta-learning-based method for cross-domain small sample population counting includes the following steps;

[0066] Step S1: Sample the synthetic crowd counting dataset and the real-world crowd counting dataset to construct the meta-training set and the meta-test set respectively;

[0067] Step S2: Construct a scene adaptation module that dynamically adjusts the network based on scene context information;

[0068] Step S3: Construct a crowd counting model;

[0069] Step S4: Design a training strategy based on meta-learning;

[0070] Step S5: Using the model obtained in step S4, test it on the meta-test set using a regularization fine-tuning method based on knowledge distillation.

[0071] Step S1 specifically includes the following steps;

[0072] Step S11: First, divide the synthesized crowd counting dataset according to the scenario. Randomly select a scenario and randomly extract several data points in the scenario as the support set and query set to form a small sample task. The support set is used for model learning, and the query set is used to evaluate the learning effect of the model.

[0073] Repeat this step to form the meta-training set D train The meta-training set is composed of units of the small sample task P in step S11, i.e. Where T train It is the number of small sample tasks in the meta-training set; and wherein is the support set of the i-th small sample task in the meta-training set, is the query set of the i-th small sample task in the meta-training set;

[0074] Step S12: first randomly select a scene from the real-world crowd counting dataset, sample data in the scene to form a support set and a query set, and form a small sample task, wherein the support set is used for model learning, and the query set is used for testing the effect of the model; repeat the step to form a meta-test set D test ; the meta-test set is formed in units of small sample tasks P in step S12, that is wherein T test is the number of small sample tasks in the meta-test set; and wherein is the support set of the i-th small sample task in the meta-test set, is the query set of the i-th small sample task in the meta-test set.

[0075] Step S2 specifically includes the following steps:

[0076] Step S21: input the feature F in from the pre-convolutional layer of the scene adaptive module into the scene adaptive module, that is, input the c x h x w feature obtained from the pre-convolutional layer into the scene adapter, and the specific expression is:

[0077] λ = softmax(w2(w1(Avg(F in ))+b1)+b2),

[0078] wherein λ represents the coefficient of the convolution kernel template obtained by the scene adapter, and the dimension size is 4, Avg(·) represents the average pooling layer, w i and b i are the weights and biases of the i-th full connection layer, and softmax(·) represents the softmax layer; step S22: multiply the coefficient obtained in step S21 by the convolution kernel template to obtain the scene adaptive convolution, and the specific expression is:

[0079]

[0080] wherein Θ d represents the scene adaptive convolution, V n represents the n-th convolution kernel template, and λ n represents the coefficient corresponding to the n-th convolution kernel template. The dimension of the convolution kernel template is C in x 1 x k x k, C in x C out x 1 x 1, and C in x Cout ×k×1 and C in ×C out ×1×k; wherein C in , C out and k represent the input channel number, output channel number and convolution kernel size of the original convolution layer at the embedding position of the scene adaptive module, respectively;

[0081] Step S23: using the scene adaptive convolution obtained in step S22 to extract scene-specific features, which are used to enhance the features extracted by the original convolution layer at the embedding position, to obtain enhanced features F out , and the specific expression is:

[0082] F out = Θ (F in ) + Θ d (F in ),

[0083] wherein Θ is the original convolution layer.

[0084] In step S3, the crowd counting model includes a feature extraction network, a density map regression network and a scene adaptive module.

[0085] Step S3 specifically includes the following steps:

[0086] Step S31: first use the first 10 layers of VGG-16 as the feature extraction network to extract high-level features of the current image;

[0087] Step S32: input the high-level features obtained in step S31 into the density map regression network, which is composed of convolution layers, and embed the scene adaptive module in a residual manner on each convolution layer.

[0088] In step S4, the training strategy is divided into two stages of pre-training and meta-training.

[0089] Step S4 specifically includes the following steps:

[0090] Step S41: train the crowd counting model constructed in step S3 on the synthetic crowd counting dataset in batches with mean square error as the loss function until the loss value converges and tends to be stable, to obtain an initial model;

[0091] Step S42: train on the meta-training set D train constructed in step S11; first fix the parameters of the feature extraction network of the initial model obtained in step S41, and when training on the i-th small sample task of D train , first perform an inner loop, that is, use the learning rate α on the support set to obtain the task-specific parameters from the model parameters θ through multiple gradient descents The specific expression is:

[0092]

[0093] Wherein L is the mean square error loss, the formula is:

[0094] L(x; θ) = ||f θ (x) - y|| 2 ,

[0095] Wherein f θ (·) refers to the crowd counting model with parameters θ, x is the input crowd image, and y is the density map label corresponding to the crowd image;

[0096] Step S43: using the task-specific parameters The outer loop is executed; the loss is calculated on the query set , and the model parameters θ are updated using the learning rate β, and the specific expression is:

[0097]

[0098] Step S44: repeat steps S42 to S43 in units of small sample tasks until all small sample tasks in the meta-training set are traversed.

[0099] In step S42, the task-specific parameters are obtained from the model parameters θ by 10 times of gradient descent

[0100] Step S5 specifically includes the following steps:

[0101] Step S51: when the i-th task test of the test meta-test set D , first create a student model and a teacher model by copying the model obtained in step S4; the student model is updated by the supervision branch and the distillation branch; the loss function is composed of the supervision loss and the distillation loss, and the specific expression is:

[0102] L T = L S + ρL D ,

[0103] Wherein ρ is the weight of the distillation loss, L S is the supervision loss, and L D is the distillation loss; on the supervision branch, the density map label of the support set is used to calculate the supervision loss L S on the support set, and the specific expression is:

[0104]

[0105] wherein θ represents a parameter S student model; on the distillation branch, the support set is input into the teacher model to generate a density map result as a pseudo label Then, the student model is supervised using the pseudo label to obtain a distillation loss L D , and the specific expression is as follows:

[0106]

[0107] Step S52: updating the teacher model by performing exponential moving average on the parameters of the student model, and the specific expression is as follows:

[0108] θ t = ηθ t + (1-η)θS ,

[0109] wherein η is a decay factor;

[0110] Step S53: repeating steps S51 to S52, and obtaining task specific parameters after performing multiple times, for testing on the query set .

[0111] Step S54: repeating steps S51 to S53 until all small sample tasks on the meta test set are traversed, and averaging the test results of all small sample tasks to obtain the evaluation result of the final model.

[0112] In step S53, the task specific parameters are obtained after repeating steps S51 to S52 twenty times.

[0113] The above program design scheme provided by the embodiment can be stored in a computer readable storage medium in a coded form, and is implemented in the form of a computer program, and the basic parameter information required for calculation is input through computer hardware, and the calculation result is output.

[0114] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, an apparatus, or a computer program product. Therefore, the present application can adopt a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.

[0115] This invention is described with reference to flowchart illustrations of methods, apparatus (devices), and computer program products according to embodiments of the invention. It will be understood that each step in the flowchart, and combinations of steps in the flowchart, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, generate instructions for implementing the steps in the flowchart. Figure 1 A device for a function specified in one or more processes.

[0116] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 A function specified in one or more flowcharts.

[0117] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 Steps of a specified function in one or more processes.

[0118] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

[0119] This patent is not limited to the above-described preferred embodiments. Anyone can derive other forms of general category multi-target tracking analysis methods based on prototype learning under the guidance of this patent. All equivalent changes and modifications made within the scope of this patent application shall fall within the scope of this patent.

Claims

1. A cross-domain small sample crowd counting method based on meta-learning, characterized in that: Comprising the following steps; Step S1: sample the synthetic crowd counting dataset and the real-world crowd counting dataset respectively to construct a meta-training set and a meta-testing set; Step S2: construct a scene adaptive module for dynamically adjusting the network according to the scene context information; Step S3: construct a crowd counting model; Step S4: design a training strategy based on meta-learning; Step S5: use the model obtained in step S4 to test on the meta-testing set by using a knowledge distillation-based regularization fine-tuning method; Step S2 specifically comprises the following steps: Step S21: inputting the feature F from the pre-convolutional layer of the scene adaptive module into the scene adaptive module in As the input of the scene adaptive module, the c x h x w feature obtained from the pre-convolutional layer is first input into the scene adaptive module, and the specific expression is: l = softmax(w2(w1(Avg(F in ))+ b1) + b2), where λ represents the coefficients of the convolution kernel template obtained by the scene adapter, the dimension size is 4, Avg(·) represents an average pooling layer, w i and b i are the weights and biases of the i-th fully connected layer, and softmax(·) represents a softmax layer. Step S22: multiply the coefficient obtained in step S21 by the convolution kernel template to obtain a scene adaptive convolution, and the specific expression is: wherein Θ d represents a scene-adaptive convolution, V n represents an nth convolution kernel template, λ n represents a coefficient corresponding to the nth convolution kernel template; the dimensions of the convolution kernel template are C in × 1 × k × k, C in × C out × 1 × 1, C in × C out × k × 1, and C in × C out × 1 × k, respectively; wherein C in , C out , and k represent the input channel number, the output channel number, and the convolution kernel size of the original convolution layer in which the scene-adaptive module is embedded, respectively. Step S23: using the scene-adaptive convolution obtained in step S22 to extract scene-specific features, which are used to enhance the features extracted by the original convolution layer of the embedded position, to obtain enhanced features F out The specific expression is: F out = Θ(F in )+ Θ d (F n ), Wherein Θ is the original convolution layer.

2. The meta-learning based cross-domain few-shot crowd counting method according to claim 1, characterized in that: Step S1 specifically Comprises the following steps; Step S11: first divide the synthetic crowd counting dataset according to the scene, randomly select a scene, and randomly sample a plurality of data in the scene as a support set and a query set to form a small sample task, wherein the support set is used for model learning, and the query set is used for evaluating the learning effect of the model; The step group of repeating this step constitutes the meta-training set D train The meta-training set is composed of small sample tasks P in step S11, that is, Where T train is the number of small sample tasks in the meta-training set; and Where is the support set of the i-th small sample task in the meta-training set, is the query set of the i-th small sample task in the meta-training set; Step S12: first randomly select a scene from the real-world crowd counting dataset, sample data in the scene to form a support set and a query set, and form a small sample task, wherein the support set is used for model learning, and the query set is used for testing the effect of the model; Repeat this step to form the meta-test set D test ; the meta-test set is formed in units of small-sample tasks P in step S12, i.e. where T test is the number of small-sample tasks in the meta-test set; and where is the support set of the i-th small-sample task in the meta-test set, is the query set of the i-th small-sample task in the meta-test set.

3. The meta-learning based cross-domain few-shot crowd counting method of claim 1, wherein: In step S3, the crowd counting model comprises a feature extraction network, a density map regression network and a scene adaptive module.

4. The meta-learning based cross-domain few-shot crowd counting method of claim 3, wherein: Step S3 specifically comprises the following steps: Step S31: first use the first 10 layers of VGG-16 as a feature extraction network to extract high-level features of the current image; Step S32: input the high-level features obtained in step S31 into the density map regression network, and the density map regression network is composed of convolution layers, and the scene adaptive module is embedded on each convolution layer in a residual manner.

5. The meta-learning based cross-domain few-shot crowd counting method of claim 2, wherein: In step S4, the training strategy is divided into two stages of pre-training and meta-training.

6. The meta-learning based cross-domain few-shot crowd counting method according to claim 5, characterized in that: Step S4 Specifically comprises the following steps: Step S41: train the crowd counting model constructed in step S3 on the synthetic crowd counting dataset in batches, and use mean square error as the loss function, until the loss value converges and tends to be stable, to obtain an initial model; Step S42: In the meta-training set D constructed by step S11 train training on the above. First, the parameters of the feature extraction network of the initial model obtained in step S41 are fixed, and the parameters of the classification network are trained on the D train The inner loop is first performed when training on the ith small sample task, that is, the support set The learning rate a is used on the support set, and the task-specific parameters of the model are obtained from the model parameters θ through multiple gradient descent The specific expression is: Wherein L is the mean square error loss, and the formula is: L(x; θ) = ||f θ (x) - y 2 , where f θ (·) denotes the crowd counting model for the parameter θ, x is the input crowd image, and y is the density map label corresponding to the crowd image. Step S43: using the task-specific parameters obtained from step S42 performing outer loop; in the query set Calculate the loss on the upper calculation, update the model parameters θ with learning rate β, the specific expression is: Step S44: repeat steps S42 to S43 in small sample task units until all small sample tasks in the meta-training set are traversed.

7. The meta-learning based cross-domain few-shot crowd counting method of claim 6, wherein: In step S42, the task-specific parameters are obtained from the model parameters θ by 10 gradient descents 8. The meta-learning based cross-domain few-shot crowd counting method of claim 2, wherein: Step S5 specifically comprises the following steps: Step S51: In the test meta test set D test The i-th task First, the student model and teacher model are created by copying the model obtained in step S4; the student model is updated by the supervised branch and the distillation branch; the loss function consists of the supervised loss and the distillation loss, and the specific expression is as follows: L T = L S + pL D , where p is the weight of distillation loss, L S is the supervision loss, L D is the distillation loss; on the supervision branch, the support set density map labels computes the supervision loss L on the support set S , specifically expressed as: Wherein The representative parameter is θ S The student model; on the distillation branch, the support set The density map result generated after the image of the support set is input to the teacher model is used as a pseudo label Then the student model is supervised using the pseudo label to obtain a distillation loss L D The specific expression is: Step S52: update the teacher model by performing exponential moving average on the parameters of the student model, and the specific expression is: θ t = ηθ t + (1 - η)θ S , Wherein η is the decay factor; Step S53: Repeat steps S51 to S52, and obtain the task after performing multiple times particular parameters for testing on the query set above; Step S54: repeat steps S51 to S53 until all small sample tasks in the meta-testing set are traversed, and average the test results of all small sample tasks to obtain the evaluation result of the final model.

9. The meta-learning-based cross-domain small sample crowd counting method according to claim 8, characterized in that: In step S53, the task is obtained after repeating steps S51 to S52 twenty times The specific parameters.

Citation Information

Patent Citations

  • Self-supervised self-distillation element learning method for small sample remote sensing image scene classification

    CN115100543A

  • View angle self-adaption-based crowd counting method in complex scene

    CN116189087A