Object statistical method, model training method, device, equipment and storage medium

The self-attention mechanism encodes and decoding networks to process image features and generates density matrix, which solves the problem of insufficient statistical accuracy in dense object counting, and achieves higher counting accuracy and efficiency.

CN115035477BActive Publication Date: 2025-07-22SHANGHAI SENSETIME TECH DEV CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210764517.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-29
Publication Date
2025-07-22
Estimated Expiration
2042-06-29

AI Technical Summary

Technical Problem

Existing deep learning algorithms have the problem of insufficient statistical accuracy in dense object counting tasks, especially inadequate feature attention under different examples, resulting in inaccurate counting results.

Method used

The example feature vector is encoded through the self-attention mechanism, a first feature sequence is generated, and combined with the image feature map to decode it to generate a density matrix to pay attention to the shared features of the object to be counted under different examples, and improve the accuracy of the statistical results.

Benefits of technology

Improves the statistical accuracy of the dense object counting task, can accurately count the number and position of objects under different examples, expands the scope of application, reduces the amount of calculation and saves box selection steps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115035477B_ABST
    Figure CN115035477B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose an object statistics method, a model training method, a device, a device, a storage medium, and a program product. Among them, the object statistics method includes: obtaining an image feature map corresponding to an image to be processed and at least one example feature vector corresponding to an object to be counted; the image to be processed includes a plurality of the objects to be counted; encoding the at least one example feature vector to obtain a first feature sequence; decoding the first feature sequence and the image feature map to obtain a density matrix; generating a statistical result for the object to be counted in the image to be processed based on the density matrix.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to, but is not limited to, the field of data processing technologies, and in particular, to an object counting method, a model training method, an apparatus, a device, and a storage medium. Background Art

[0002] In recent years, deep learning algorithms have made great progress in various fields and have also been applied to many scenario analysis fields. In the field of scenario analysis, counting dense objects is a very important task. Currently, there are many fully automated deep learning-based image-oriented counting methods, such as counting people, counting goats, counting airport luggage, etc. These methods use deep learning models to use image information to perform fast and automated counting of specific objects, and have been widely applied in scenarios such as people counting in the monitoring field, animal counting in nature conservation applications, and workpiece counting in factory industrial manufacturing. Summary of the Invention

[0003] In view of this, embodiments of the present disclosure at least provide an object counting method, a model training method, an apparatus, a device, a storage medium, and a program product.

[0004] The technical solution of the embodiments of the present disclosure is implemented as follows:

[0005] On the one hand, embodiments of the present disclosure provide an object counting method, the method including: obtaining an image feature map corresponding to an image to be processed and at least one example feature vector corresponding to an object to be counted; the image to be processed includes a plurality of the objects to be counted; performing encoding processing on the at least one example feature vector to obtain a first feature sequence; performing decoding processing on the first feature sequence and the image feature map to obtain a density matrix; generating a counting result for the object to be counted in the image to be processed based on the density matrix.

[0006] In some embodiments, the performing decoding processing on the first feature sequence and the image feature map to obtain a density matrix includes: performing first decoding processing on the image feature map to obtain a second feature sequence; performing second decoding processing on the first feature sequence and the second feature sequence to obtain the density matrix.

[0007] Through the embodiments of the present disclosure, since the first feature sequence representing the common features of the object to be counted under different examples decodes the second feature sequence corresponding to the image feature map, the obtained density matrix can, on the basis of the image feature map, focus on the common features of the object to be counted under different examples, providing a data basis for subsequent counting tasks and indirectly improving the accuracy of the counting result.

[0008] In some embodiments, the second decoding process includes multiple iterative sub - decoding processes; the second decoding process of the first feature sequence and the second feature sequence to obtain the density matrix includes: taking the second feature sequence as the input feature corresponding to the first sub - decoding process in the multiple sub - decoding processes; for each sub - decoding process, taking the first feature sequence and the input feature corresponding to the sub - decoding process as the input of the sub - decoding process to obtain a third feature sequence output by the sub - decoding process, and taking the third feature sequence output by the sub - decoding process as the input feature corresponding to the next sub - decoding process; wherein, the density matrix is determined based on the third feature sequence output by the last sub - decoding process in the multiple sub - decoding processes.

[0009] Through the embodiments of the present disclosure, due to the multiple iterative sub - decoding processes of the first feature sequence and the second feature sequence, in each sub - decoding process, the second feature sequence can be iteratively sub - decoded by using the first feature sequence representing the common features of the object to be statistically analyzed under different examples, which can improve the accuracy of statistics.

[0010] In some embodiments, the encoding process includes multiple iterative sub - encoding processes; the encoding process of the at least one example feature vector to obtain a first feature sequence includes: taking the at least one example feature vector as the input feature corresponding to the first sub - encoding process in the multiple sub - encoding processes; for each sub - encoding process, taking the input feature corresponding to the sub - encoding process as the input of the sub - encoding process to obtain a fourth feature sequence output by the sub - encoding process, and taking the fourth feature sequence output by the sub - encoding process as the input feature corresponding to the next sub - encoding process; wherein, the first feature sequence is the fourth feature sequence output by the last sub - encoding process in the multiple sub - encoding processes.

[0011] Through the embodiments of the present disclosure, due to the multiple iterative sub - encoding processes of the at least one example feature vector, the generalization features corresponding to the object to be statistically analyzed in each example feature vector can be strengthened, and other features corresponding to the object to be statistically analyzed in each example feature vector can be weakened. Furthermore, the first feature sequence can represent the common features of the object to be statistically analyzed under different postures (different examples), which can improve the accuracy of subsequent statistics.

[0012] In some embodiments, obtaining the image feature map corresponding to the image to be processed and at least one example feature vector corresponding to the object to be counted includes: obtaining the image to be processed and at least one detection box corresponding to the image to be processed; the detection box is used to determine the area of the object to be counted in the image to be processed; extracting the image feature map corresponding to the image to be processed; and determining the example feature vector corresponding to each detection box based on at least one detection box and the image feature map.

[0013] Through the embodiments of the present disclosure, since at least one detection box corresponding to the object to be counted is obtained, after the image feature map corresponding to the image to be processed is obtained by feature extraction on the image to be processed, the example feature vector corresponding to each detection box can be directly intercepted from the image feature map based on the detection box, thereby saving the process of extracting the example feature vector, reducing the amount of calculation, and improving the statistical efficiency.

[0014] In some embodiments, extracting the image feature map corresponding to the image to be processed includes: performing feature extraction processing on the image to be processed at multiple scales to obtain an intermediate feature map corresponding to each scale; the sizes of the intermediate feature maps at different scales are different; and fusing the intermediate feature maps corresponding to each scale to obtain the image feature map.

[0015] Through the embodiments of the present disclosure, since the image feature map obtained by performing feature extraction processing on the image to be processed at multiple scales can not only focus on the object to be counted with a larger size in the image to be processed, but also focus on the object to be counted with a smaller size in the image to be processed, the statistical accuracy of the object to be counted in the image to be processed is improved.

[0016] In some embodiments, determining the example feature vector corresponding to each detection box based on at least one detection box and the image feature map includes: obtaining the size ratio between the image feature map and the image to be processed; performing scale transformation on each detection box based on the size ratio to obtain at least one feature interception box; determining the example feature map corresponding to each detection box based on the at least one feature interception box and the image feature map; and determining the example feature vector corresponding to each detection box based on the example feature map corresponding to each detection box.

[0017] Through the embodiments of the present disclosure, since each detection box is subjected to scale transformation based on the size ratio between the image feature map and the image to be processed, and then the example feature map corresponding to each detection box is intercepted, and then the corresponding example feature vector is obtained, each example feature vector can only represent the feature information of the object to be counted within the current detection box, and the first feature sequence based on multiple example feature vectors can represent the common features of the object to be counted under different examples.

[0018] In some embodiments, obtaining the image feature map corresponding to the image to be processed and at least one example feature vector corresponding to the object to be counted includes: obtaining the image to be processed and at least one example image; the example image includes the object to be counted, and the example image is different from the image to be processed; extracting the image feature map corresponding to the image to be processed; and extracting the example feature vector corresponding to each example image.

[0019] Through the embodiments of the present disclosure, since feature extraction processing is performed on the image to be processed and at least one example image respectively, in the actual application process, it is not limited to selecting the example image in the image to be processed. That is to say, the example image can be obtained from other images that are not the image to be processed, which improves the application scope of the embodiments of the present disclosure; at the same time, for the application scenario of a large number of images to be processed, the embodiments of the present disclosure only need to determine the example image once, and then can directly perform the counting process for the same object to be counted on each image to be processed, saving the selection step for each image to be processed.

[0020] In some embodiments, generating the statistical result for the object to be counted in the image to be processed based on the density matrix includes: generating a density estimation map based on the density matrix; and generating the statistical result for the object to be counted in the image to be processed based on the density estimation map.

[0021] In some embodiments, generating a density estimation map based on the density matrix includes: performing iterative convolution processing on the density matrix; resampling the density matrix after iterative convolution processing to obtain a density estimation map with the same size as the image to be processed; generating the statistical result for the object to be counted in the image to be processed based on the density estimation map includes: summing the density estimation map to obtain the statistical quantity of the object to be counted in the image to be processed; and / or, obtaining local peak points in the density estimation map, and screening the local peak points through a non-maximum suppression algorithm to obtain the position of the object to be counted; the pixel value of the local peak point is greater than the pixel values of adjacent pixel points.

[0022] Through the above embodiments, not only can the statistical quantity of the object to be counted in the image to be processed be obtained, but also the position of each object to be counted in the image to be processed can be obtained, expanding the application scope of the present disclosure.

[0023] On the other hand, the embodiments of the present disclosure provide a method for training an object counting model, and the method includes:

[0024] Obtain an image feature map corresponding to the sample image and at least one example feature vector corresponding to the object to be counted based on the feature acquisition network; the sample image includes a plurality of the objects to be counted; perform encoding processing on the at least one example feature map through the encoding and decoding network to obtain a first feature sequence; perform decoding processing on the first feature sequence and the image feature map to obtain the density matrix; convert the density matrix into a predicted density map through the density regression network; adjust the parameters of the object counting model to be trained based on the predicted density map and the true density map corresponding to the sample image to obtain a trained object counting model.

[0025] In some embodiments, the method further includes: constructing an initial matrix; the size of the initial matrix is the same as the size of the sample image; obtaining annotation information for the sample image; the annotation information is used to represent the relative coordinates of each object to be counted in the sample image; update the element value of the matrix position corresponding to the initial matrix based on the relative coordinates of each object to be counted in the sample image to obtain the true density map.

[0026] In some embodiments, the encoding and decoding network includes an encoding network and a decoding network; the performing encoding processing on the at least one example feature map through the encoding and decoding network to obtain a first feature sequence; performing decoding processing on the first feature sequence and the image feature map to obtain the density matrix includes: performing encoding processing on the at least one example feature vector through the encoding network to obtain a first feature sequence; performing decoding processing on the first feature sequence and the image feature map through the decoding network to obtain the density matrix.

[0027] In some embodiments, the decoding network includes a first decoding layer and a second decoding layer; the performing decoding processing on the first feature sequence and the image feature map through the decoding network to obtain the density matrix includes: performing first decoding processing on the image feature map through the first decoding layer to obtain a second feature sequence; performing second decoding processing on the first feature sequence and the second feature sequence through the second decoding layer to obtain the density matrix.

[0028] In some embodiments, the second decoding layer includes a plurality of cascaded sub-decoding layers; the second decoding process of the first feature sequence and the second feature sequence through the second decoding layer to obtain the density matrix includes: taking the second feature sequence as the input feature corresponding to the first sub-decoding layer among the plurality of sub-decoding layers; for each sub-decoding layer, inputting the first feature sequence and the input feature corresponding to the sub-decoding layer into the sub-decoding layer to obtain a third feature sequence output by the sub-decoding layer, and taking the third feature sequence output by the sub-decoding layer as the input feature corresponding to the next sub-decoding layer; wherein, the density matrix is determined based on the third feature sequence output by the last sub-decoding layer among the plurality of sub-decoding layers.

[0029] In some embodiments, the encoding layer includes a plurality of cascaded sub-encoding layers; the encoding process of the at least one example feature vector through the encoding network to obtain a first feature sequence includes: taking the at least one example feature vector as the input feature corresponding to the first sub-encoding layer among the plurality of sub-encoding layers; for each sub-encoding layer, inputting the input feature corresponding to the sub-encoding layer into the sub-encoding layer to obtain a fourth feature sequence output by the sub-encoding layer, and taking the fourth feature sequence output by the sub-encoding layer as the input feature corresponding to the next sub-encoding layer; wherein, the first feature sequence is the fourth feature sequence output by the last sub-encoding layer among the plurality of sub-encoding layers.

[0030] In some embodiments, the feature acquisition network includes a feature extraction network and a feature processing network; the acquisition of the image feature map corresponding to the sample image and at least one example feature vector corresponding to the object to be counted based on the feature acquisition network includes: acquiring the sample image and at least one detection frame corresponding to the sample image; the detection frame is used to determine the area of the object to be counted in the sample image; extracting the image feature map corresponding to the sample image through the feature extraction network; based on the feature processing network, determining the example feature vector corresponding to each detection frame through at least one detection frame and the image feature map.

[0031] In some embodiments, the feature extraction network includes a feature fusion layer and a plurality of cascaded first convolutional layers; the extraction of the image feature map corresponding to the sample image through the feature extraction network includes: performing feature extraction processing on the sample image through the plurality of cascaded first convolutional layers to obtain an intermediate feature map output by each first convolutional layer; the sizes of the intermediate feature maps output by different first convolutional layers are different; fusing the intermediate feature maps output by each first convolutional layer through the feature fusion layer to obtain the image feature map.

[0032] In some embodiments, the feature processing network includes a feature truncation layer and a pooling layer; based on the feature processing network, determining an example feature vector corresponding to each detection box through at least one detection box and the image feature map includes: obtaining the size ratio between the image feature map and the sample image through the feature truncation layer, performing scale transformation on each detection box based on the size ratio to obtain at least one feature truncation box; determining an example feature map corresponding to each detection box based on the at least one feature truncation box and the image feature map; and converting the example feature map corresponding to each detection box through the pooling layer to obtain an example feature vector corresponding to each detection box.

[0033] In some embodiments, the feature acquisition network includes a feature extraction network, a pooling layer, a feature fusion layer, and a plurality of cascaded first convolutional layers; based on the feature acquisition network, obtaining an image feature map corresponding to a sample image and at least one example feature vector corresponding to an object to be counted includes: obtaining the sample image and at least one example image; the example image includes the object to be counted, and the example image is different from the sample image; performing feature extraction processing on the sample image through the feature extraction layer to obtain an image feature map corresponding to the sample image; performing feature extraction processing on at least one example image through the feature extraction layer to obtain at least one example feature map; and converting the at least one example feature map into at least one example feature vector through the pooling layer.

[0034] On the other hand, an object counting device provided by an embodiment of the present disclosure includes:

[0035] A first acquisition module, configured to acquire an image feature map corresponding to a to-be-processed image and at least one example feature vector corresponding to an object to be counted; the to-be-processed image includes a plurality of the objects to be counted;

[0036] An encoding module, configured to perform encoding processing on the at least one example feature vector to obtain a first feature sequence;

[0037] A decoding module, configured to perform decoding processing on the first feature sequence and the image feature map to obtain a density matrix;

[0038] A generation module, configured to generate a statistical result for the object to be counted in the to-be-processed image based on the density matrix.

[0039] On the other hand, an object counting model training device provided by an embodiment of the present disclosure, the object counting model includes a feature acquisition network, an encoding and decoding network, and a density regression network, and the device includes:

[0040] A second acquisition module, configured to obtain an image feature map corresponding to a sample image and at least one example feature vector corresponding to an object to be counted based on the feature acquisition network; the sample image includes a plurality of the objects to be counted;

[0041] An encoding and decoding module, configured to perform encoding processing on the at least one example feature map through the encoding and decoding network to obtain a first feature sequence; perform decoding processing on the first feature sequence and the image feature map to obtain the density matrix;

[0042] A regression module, configured to convert the density matrix into a predicted density map through the density regression network;

[0043] An adjustment module, configured to adjust parameters of the object counting model to be trained based on the predicted density map and the true density map corresponding to the sample image, so as to obtain a trained object counting model.

[0044] In another aspect, an embodiment of the present disclosure provides a computer device, including a memory and a processor, where the memory stores a computer program that can run on the processor, and when the processor executes the program, it implements some or all of the steps in the above method.

[0045] In another aspect, an embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements some or all of the steps in the above method.

[0046] In another aspect, an embodiment of the present disclosure provides a computer program, including computer-readable code, and when the computer-readable code runs in a computer device, a processor in the computer device executes to implement some or all of the steps in the above method.

[0047] In another aspect, an embodiment of the present disclosure provides a computer program product, where the computer program product includes a non-transitory computer-readable storage medium storing a computer program, and when the computer program is read and executed by a computer, it implements some or all of the steps in the above method.

[0048] In the embodiments of the present disclosure, by performing encoding processing on the at least one example feature vector through the self-attention mechanism to obtain a first feature sequence, common features of the object to be counted under different examples can be focused on the basis of each example feature vector; at the same time, since the density matrix is obtained by performing decoding processing on the first feature sequence and the image feature map, in this way, the obtained density matrix can focus on the common features of the object to be counted under different examples on the basis of the image feature map, so that the accuracy of the statistical result can be improved in the process of generating a statistical result for the object to be counted based on the density matrix.

[0049] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the technical solutions of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The drawings herein are incorporated into the specification and constitute a part of this specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.

[0051] Figure 1 Schematic diagram of the implementation process of an object counting method provided by an embodiment of the present disclosure;

[0052] Figure 2 Schematic diagram of the implementation process of an object counting method provided by an embodiment of the present disclosure;

[0053] Figure 3 Schematic diagram of the implementation process of an object counting method provided by an embodiment of the present disclosure;

[0054] Figure 4 Schematic diagram of the implementation process of an object counting method provided by an embodiment of the present disclosure;

[0055] Figure 5 Schematic diagram of the implementation process of an object counting method provided by an embodiment of the present disclosure;

[0056] Figure 6 Schematic diagram of the implementation process of a training method for an object counting model provided by an embodiment of the present disclosure;

[0057] Figure 7 Schematic diagram of the implementation process of a training method for an object counting model provided by an embodiment of the present disclosure;

[0058] Figure 8a Schematic diagram of a real image provided by an embodiment of the present disclosure;

[0059] Figure 8b Schematic diagram of a real image with a rectangular bounding box added provided by an embodiment of the present disclosure;

[0060] Figure 8c Schematic diagram of a real density map provided by an embodiment of the present disclosure;

[0061] Figure 8d Schematic diagram of a network structure module provided by an embodiment of the present disclosure;

[0062] Figure 9 Schematic diagram of the composition structure of an object counting device provided by an embodiment of the present disclosure;

[0063] Figure 10Schematic diagram of the composition structure of a training device for an object statistical model provided by an embodiment of the present disclosure;

[0064] Figure 11 Schematic diagram of the hardware entity of a computer device provided by an embodiment of the present disclosure. Detailed implementation manners

[0065] In order to make the objectives, technical solutions, and advantages of the present disclosure clearer, the technical solutions of the present disclosure will be further elaborated in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be construed as limitations on the present disclosure. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present disclosure.

[0066] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict. The terms "first / second / third" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein.

[0067] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs. The terms used herein are only for the purpose of describing the present disclosure and are not intended to limit the present disclosure.

[0068] An embodiment of the present disclosure provides an object statistical method, which can be executed by a processor of a computer device. Among them, the computer device may refer to a device with data processing capabilities such as a server, a laptop computer, a tablet computer, a desktop computer, a smart TV, a set-top box, a mobile device (such as a mobile phone, a portable video player, a personal digital assistant, a dedicated messaging device, a portable game device), etc.

[0069] Figure 1 Schematic diagram of the implementation process of an object statistical method provided by an embodiment of the present disclosure, as Figure 1 shown, the method includes the following steps S101 to step S103:

[0070] Step S101, obtaining an image feature map corresponding to the image to be processed and at least one example feature vector corresponding to the object to be counted; the image to be processed includes a plurality of the objects to be counted.

[0071] Among them, the image to be processed is an image for which the quantity of the object to be counted needs to be counted. In some implementation scenarios, the image to be processed includes a large number of objects to be counted, and the object counting method provided by the embodiments of the present disclosure is used to count the objects to be counted in the image to be processed. In some embodiments, the image features of the image to be processed can be extracted through a pre-trained feature extraction network to obtain an image feature map corresponding to the image to be processed.

[0072] In some embodiments, each example feature vector corresponding to the object to be counted can be obtained based on an example image including the object to be counted. The image features of each example image can be extracted respectively by using the pre-trained feature extraction network to obtain at least one example feature vector.

[0073] Among them, the example image corresponding to the object to be counted is an image including the object to be counted, and the at least one example image includes the same object to be counted. That is, the object to be counted that needs to be counted can be determined through the at least one example image. That is to say, the image to be processed can include multiple different types of objects. When the objects to be counted determined by the input at least one example image are different, the embodiments of the present disclosure can output the statistical results corresponding to different objects to be counted.

[0074] In some embodiments, the example image can be from the same image as the image to be processed. In this embodiment, the user can select at least one object to be counted in the image to be processed to obtain at least one detection box of the object to be counted. Through the at least one detection box corresponding to the object to be counted, the corresponding example image can be intercepted from the example image.

[0075] In some embodiments, the example image can be from a different image from the image to be processed. Exemplarily, when it is necessary to count the object to be processed in image A, the user can select at least one object to be counted in image B to obtain at least one detection box of the object to be counted, and then obtain an example image from image B; the user can also select at least one object to be counted in image C to obtain at least one detection box of the object to be counted, and then obtain an example image from image C; the example images from image B and image C are used as the at least one example image corresponding to the object to be counted.

[0076] It should be noted that there is at least one object to be counted in the example image. That is to say, in the example image, there can be one object to be counted, or two or more objects to be counted.

[0077] Step S102: Perform encoding processing on the at least one example feature vector to obtain a first feature sequence.

[0078] In some embodiments, the at least one example feature vector can be encoded through a self-attention mechanism to obtain a first feature sequence. During the process of encoding the at least one example feature vector based on the self-attention mechanism, the at least one example feature vector can be used as the query feature, value feature, and key feature of the self-attention mechanism to obtain the first feature sequence output by the self-attention mechanism.

[0079] Among them, since the at least one example feature vector characterizes the feature information of the object to be counted in different example images, the first feature sequence obtained based on the self-attention mechanism can focus on the common features of the object to be counted under different examples on the basis of each example feature vector.

[0080] Step S103: Decode the first feature sequence and the image feature map to obtain a density matrix.

[0081] In the embodiments of the present invention, the above steps of encoding the at least one example feature vector to obtain a first feature sequence and decoding the first feature sequence and the image feature map to obtain a density matrix can be completed through an encoder-decoder network. The encoder-decoder network includes an encoder network and a decoder network, where the encoder network and the decoder network can be implemented based on their respective network models, and the algorithms used by the encoder network and the decoder network in the encoder-decoder network are not limited. The network model can be, but is not limited to, a Convolutional Neural Networks (CNN) model, a Recurrent Neural Networks (RNN) model, a Long Short-Term Memory (LSTM), a delay network model, or a gated convolutional neural network model, etc.

[0082] Among them, the encoder network can include an encoding layer implemented based on the self-attention mechanism. For example, the encoding layer can be implemented based on a Self-Attention Network (SAN). The decoder network can include a decoding layer based on the self-attention mechanism and a decoding layer based on the multi-head attention mechanism.

[0083] Exemplarily, the encoder-decoder network can be a Transformer model. The encoder network based on the Transformer model encodes the at least one example feature vector to obtain a first feature sequence; the decoder network based on the Transformer model decodes the first feature sequence and the image feature map to obtain a density matrix.

[0084] Step S104: Generate a statistical result for the object to be counted in the image to be processed based on the density matrix.

[0085] In some embodiments, the statistical result may include at least one of the following: the statistical quantity of the object to be counted in the image to be processed, and the position of the object to be counted in the image to be processed.

[0086] In the embodiments of the present disclosure, by encoding the at least one example feature vector through a self-attention mechanism, a first feature sequence can focus on the common features of the object to be counted under different examples on the basis of each example feature vector; at the same time, since the density matrix is obtained by decoding the first feature sequence and the image feature map, in this way, the obtained density matrix can focus on the common features of the object to be counted under different examples on the basis of the image feature map, so that the accuracy of the statistical result can be improved in the process of generating the statistical result for the object to be counted based on the density matrix.

[0087] Figure 2 is an optional process schematic diagram of the object statistical method provided by the embodiments of the present disclosure, and this method can be executed by the processor of a computer device. Based on Figure 1 , Figure 1 S103 in Figure 2 can be updated to S201 to S202, and will be described in combination with the steps shown in

[0088] Step S201: Perform a first decoding process on the image feature map to obtain a second feature sequence.

[0089] In some embodiments, the first decoding process on the image feature map can be performed through a first attention mechanism to obtain a second feature sequence. Among them, the first attention mechanism is a self-attention mechanism. In the process of performing the first decoding process on the image feature map based on the self-attention mechanism, the image feature map can be used as the query feature, value feature, and key feature of the self-attention mechanism to obtain the second feature sequence output by the self-attention mechanism.

[0090] In some embodiments, the first decoding process on the image feature map can be performed through a decoding network based on an attention mechanism to obtain the second feature sequence. Among them, the decoding network can include a first decoding layer based on a self-attention mechanism, and through the first decoding layer, the first decoding process on the image feature map can be performed to obtain the second feature sequence.

[0091] Among them, the image feature map can be expressed as h×w×c. During the process of performing the first decoding process on the image feature map, it is necessary to rearrange the image feature map into an image feature sequence. Exemplarily, according to the dimension of h, the corresponding (w×c) can be spliced into a sequence to obtain the image feature sequence, and the image feature sequence can be expressed as 1×(h×w)×c. Then, the image feature sequence 1×(h×w)×c is used as the input of the first decoding layer, that is, as the query feature, value feature, and key feature of the self-attention mechanism, to obtain the second feature sequence output by the first decoding layer (self-attention mechanism). Correspondingly, the second feature sequence can also be expressed as 1×(h×w)×c.

[0092] Step S202: Perform a second decoding process on the first feature sequence and the second feature sequence to obtain the density matrix.

[0093] In some embodiments, the first feature sequence and the second feature sequence can be subjected to a second decoding process through a second attention mechanism to obtain the density matrix. Among them, the second attention mechanism is a multi-head attention mechanism. During the process of performing the second decoding process on the first feature sequence and the second feature sequence based on the multi-head attention mechanism, the first feature sequence can be used as the value feature and key feature of the multi-head attention mechanism, and the second feature sequence can be used as the query feature of the multi-head attention mechanism to obtain the density matrix output by the multi-head attention mechanism.

[0094] In some embodiments, the first feature sequence and the second feature sequence can be subjected to a second decoding process through the above decoding network based on the multi-head attention mechanism to obtain the density matrix. Among them, the decoding network can include a second decoding layer based on the multi-head attention mechanism, and through the second decoding layer, the first feature sequence and the second feature sequence can be subjected to a second decoding process to obtain the density matrix.

[0095] Among them, the first feature sequence can be expressed as 1×n×c, where n is the number of example feature vectors, and the second feature sequence can be expressed as 1×(h×w)×c. During the second decoding process of the first feature sequence and the second feature sequence, the first feature sequence 1×n×c is used as the value feature and key feature of the second decoding layer (attention mechanism), and the second feature sequence 1×(h×w)×c is used as the query feature of the attention mechanism to obtain the attention result output by the second decoding layer (attention mechanism). Then, the attention result is input into the feedforward neural network included in the decoding network, and a density matrix can be obtained. The density matrix can be expressed as h×w×c, that is, a two-dimensional matrix of h×w, and each matrix position of the two-dimensional matrix corresponds to a vector with a length of c.

[0096] In some embodiments, the second decoding process includes multiple iterative sub-decoding processes. The above-mentioned second decoding process of the first feature sequence and the second feature sequence to obtain the density matrix can be implemented through step S2021 to step S2022.

[0097] Step S2021: Use the second feature sequence as the input feature corresponding to the first sub-decoding process in the multiple sub-decoding processes.

[0098] Step S2022: For each sub-decoding process, use the first feature sequence and the input feature corresponding to the sub-decoding process as the input of the sub-decoding process to obtain the third feature sequence output by the sub-decoding process, and use the third feature sequence output by the sub-decoding process as the input feature corresponding to the next sub-decoding process.

[0099] Among them, the density matrix is determined based on the third feature sequence output by the last sub-decoding process in the multiple sub-decoding processes.

[0100] In some embodiments, for the first sub-decoding process in the multiple sub-decoding processes, the input feature corresponding to the first sub-decoding process is the second feature sequence and the first feature sequence, and the output feature of the first sub-decoding process is the third feature sequence. And the third feature sequence output by the first sub-decoding process is the input feature corresponding to the second sub-decoding process.

[0101] In some embodiments, for the Nth sub-decoding process in the multiple sub-decoding processes, the input feature corresponding to the Nth sub-decoding process is the third feature sequence output by the (N - 1)th sub-decoding process and the first feature sequence, and the output feature of the Nth sub-decoding process is the third feature sequence. And the third feature sequence output by the Nth sub-decoding process is the input feature corresponding to the (N + 1)th sub-decoding process.

[0102] In some embodiments, for the last sub-decoding process among multiple sub-decoding processes, a density matrix may be generated based on a third feature sequence output by the last sub-decoding process. Exemplarily, the third feature sequence output by the last sub-decoding process may be input into a feedforward neural network to obtain the density matrix generated by the feedforward neural network.

[0103] Through the embodiments of the present disclosure, due to performing iterative multiple sub-decoding processes on the first feature sequence and the second feature sequence, in each sub-decoding process, the second feature sequence can be iteratively sub-decoded by using the first feature sequence representing the common features of the object to be statistically analyzed under different examples, which can improve the accuracy of the statistics.

[0104] In some embodiments, in the case where the second decoding process includes iterative multiple sub-decoding processes, Figure 1 The encoding process in S102 may also be iterative multiple sub-encoding processes. In some embodiments, the number of iterations of the sub-encoding process and the sub-decoding process is the same. The above encoding process for the at least one example feature vector to obtain a first feature sequence includes: using the at least one example feature vector as the input feature corresponding to the first sub-encoding process among the multiple sub-encoding processes. For each sub-encoding process, using the input feature corresponding to the sub-encoding process as the input of the sub-encoding process to obtain a fourth feature sequence output by the sub-encoding process, and using the fourth feature sequence output by the sub-encoding process as the input feature corresponding to the next sub-encoding process. Wherein, the first feature sequence is the fourth feature sequence output by the last sub-encoding process among the multiple sub-encoding processes.

[0105] Through the embodiments of the present disclosure, due to performing iterative multiple sub-encoding processes on the at least one example feature vector, the generalization features corresponding to the object to be statistically analyzed in each example feature vector can be strengthened, and other features corresponding to the object to be statistically analyzed in each example feature vector can be weakened. Furthermore, the first feature sequence can represent the common features of the object to be statistically analyzed under different postures (different examples), which can improve the accuracy of subsequent statistics.

[0106] Through the embodiments of the present disclosure, due to decoding the second feature sequence corresponding to the image feature map by using the first feature sequence representing the common features of the object to be statistically analyzed under different examples, the obtained density matrix can, based on the image feature map, focus on the common features of the object to be statistically analyzed under different examples, provide a data basis for subsequent statistical tasks, and indirectly improve the accuracy of the statistical results.

[0107] Figure 3FIG. 0 is an alternative flowchart of the object counting method provided by the embodiments of the present disclosure, and this method can be executed by the processor of a computer device. Based on Figure 1 , Figure 1 S101 in [reference] can be updated to S301 to S303, and will be described in combination with the steps shown in Figure 3 .

[0108] Step S301: Obtain the to-be-processed image and at least one detection box corresponding to the to-be-processed image; the detection box is used to determine the area of the object to be counted in the to-be-processed image.

[0109] Step S302: Extract the image feature map corresponding to the to-be-processed image.

[0110] In some embodiments, the above-mentioned extraction of the image feature map corresponding to the to-be-processed image can be implemented through steps S3021 to S3022.

[0111] Step S3021: Perform feature extraction processing on the to-be-processed image at multiple scales to obtain intermediate feature maps corresponding to each scale; the sizes of the intermediate feature maps at different scales are different.

[0112] Step S3022: Fuse the intermediate feature maps corresponding to each scale to obtain the image feature map.

[0113] Among them, the above-mentioned fusion of the intermediate feature maps corresponding to each scale to obtain the image feature map can be implemented in the following manner: Resample the intermediate feature map corresponding to each scale to a preset size to obtain the resampled intermediate feature map corresponding to each scale; merge the resampled intermediate feature maps corresponding to each scale based on the channel dimension to obtain the image feature map with the preset size.

[0114] Through the embodiments of the present disclosure, since the to-be-processed image is subjected to feature extraction processing at multiple scales, the obtained image feature map can not only focus on the objects to be counted with larger sizes in the to-be-processed image, but also focus on the objects to be counted with smaller sizes in the to-be-processed image, improving the statistical accuracy of the objects to be counted in the to-be-processed image.

[0115] Step S303: Determine the example feature vector corresponding to each detection box based on at least one detection box and the image feature map.

[0116] In some embodiments, the above-mentioned determination of the example feature vector corresponding to each detection box based on at least one detection box and the image feature map can be implemented through steps S3031 to S3024.

[0117] Step S3031: Obtain the size ratio between the image feature map and the image to be processed.

[0118] Step S3032: Based on the size ratio, perform scale transformation on each detection box to obtain at least one feature extraction box.

[0119] In some embodiments, the size ratio may include the size ratio in the h dimension and the size ratio in the w dimension between the image feature map and the image to be processed. Based on the obtained size ratio in the h dimension and the size ratio in the w dimension, corresponding scale transformation can be performed on the coordinates of the detection box to obtain the coordinates of the feature extraction box corresponding to the detection box.

[0120] Exemplarily, if the size of the image feature map is h1×w1, the size of the image to be processed is h2×w2, and based on Step S3031, it can be determined that the size ratio between the image feature map and the image to be processed includes h1:h2 = 1:2 and w1:w2 = 1:3, then corresponding scale transformation can be performed on the vertex coordinates of the detection box based on this size ratio. That is, when the vertex coordinates of the detection box are (H, W), the vertex coordinates of the corresponding feature extraction box are (H / 2, W / 3).

[0121] Step S3033: Based on the at least one feature extraction box and the image feature map, determine the example feature map corresponding to each detection box.

[0122] In some embodiments, based on the coordinates in the h dimension and the coordinates in the w dimension of each feature extraction box, the image feature map can be intercepted to obtain the example feature map corresponding to the feature extraction box (detection box). Among them, during the interception process, the values and sizes of the example feature map in the c dimension do not change compared with those of the image feature map in the c dimension.

[0123] Step S3034: Based on the example feature map corresponding to each detection box, determine the example feature vector corresponding to each detection box.

[0124] In some embodiments, global maximum pooling can be performed on the example feature map corresponding to each detection box to obtain the example feature vector corresponding to each detection box.

[0125] Through the embodiments of the present disclosure, since scale transformation is performed on each detection box based on the size ratio between the image feature map and the image to be processed, and then the example feature map corresponding to each detection box is intercepted, and then the corresponding example feature vector is obtained, it can be ensured that each example feature vector only represents the feature information of the object to be counted within the current detection box, and the first feature sequence based on multiple example feature vectors can characterize the common features of the object to be counted under different examples.

[0126] According to the embodiments of the present disclosure, since at least one detection box corresponding to the object to be counted is obtained, after the feature extraction of the image to be processed to obtain the corresponding image feature map, the example feature vectors corresponding to each detection box can be directly intercepted in the image feature map based on the detection box, thereby saving the process of extracting the example feature vectors, reducing the amount of calculation, and improving the counting efficiency.

[0127] Figure 4 is an optional flowchart of the object counting method provided by the embodiments of the present disclosure, and this method can be executed by the processor of a computer device. Based on Figure 1 , Figure 1 S101 in can be updated to S401 to S403, and will be described in combination with Figure 4 the steps shown.

[0128] Step S401, obtain the image to be processed and at least one example image; the example image includes the object to be counted, and the example image is different from the image to be processed.

[0129] In some embodiments, the example image can be from a different image than the image to be processed. Exemplarily, when it is necessary to count the object to be processed in Image A, the user can select at least one object to be counted in Image B to obtain at least one detection box of the object to be counted, and then obtain the example image from Image B; the user can also select at least one object to be counted in Image C to obtain at least one detection box of the object to be counted, and then obtain the example image from Image C; the example images from Image B and Image C are used as at least one example image corresponding to the object to be counted.

[0130] Step S402, extract the image feature map corresponding to the image to be processed.

[0131] In some embodiments, perform feature extraction processing on the image to be processed at multiple scales to obtain intermediate feature maps corresponding to each scale; the sizes of the intermediate feature maps at different scales are different. Fuse the intermediate feature maps corresponding to each scale to obtain the image feature map.

[0132] Step S403, extract the example feature vectors corresponding to each example image.

[0133] In some embodiments, for each example image, perform feature extraction processing on the example image at multiple scales to obtain intermediate feature maps corresponding to each scale; the sizes of the intermediate feature maps at different scales are different. Fuse the intermediate feature maps corresponding to each scale to obtain the example feature vectors corresponding to the example image.

[0134] Through the embodiments of the present disclosure, since feature extraction processing is performed on the image to be processed and at least one example image separately, in the actual application process, it is not limited to selecting the example image in the image to be processed. That is to say, the example image can be obtained from other images that are not the image to be processed, which improves the application scope of the embodiments of the present disclosure. At the same time, for the application scenario of a large number of images to be processed, the embodiments of the present disclosure only need to determine the example image once, and then can directly perform the statistical process for the same object to be counted on each image to be processed, saving the selection step for each image to be processed.

[0135] Figure 5 is an optional flowchart of the object statistical method provided by the embodiments of the present disclosure, and this method can be executed by the processor of a computer device. Based on Figure 1 , Figure 1 S104 in Figure 5 can be updated to S501 to S503, and will be described in combination with the steps shown in

[0136] Step S501: Generate a density estimation map based on the density matrix.

[0137] In some embodiments, the above-mentioned generating a density estimation map based on the density matrix can be implemented through S5011 to S5012.

[0138] S5011: Perform iterative convolution processing on the density matrix.

[0139] S5012: Resample the density matrix after iterative convolution processing to obtain a density estimation map with the same size as the image to be processed.

[0140] In some embodiments, through the above resampling operation, the size of the density matrix can be adjusted to the size of the image to be processed, and the resampled density matrix is used as the density estimation map. Thus, the numerical value of each position in the density estimation map (i.e., the numerical value of the matrix element at each position in the resampled density matrix) can represent the probability that there is an object to be counted at the corresponding position in the image to be processed.

[0141] Among them, the higher the probability that there is an object to be counted at a position in the image to be processed, the larger the numerical value at that position in the density estimation map.

[0142] Step S502: Generate a statistical result for the object to be counted in the image to be processed based on the density estimation map.

[0143] In some embodiments, the above-mentioned generating a statistical result for the object to be counted in the image to be processed based on the density estimation map can be implemented through S5021 and / or S5022.

[0144] S5021, Sum the density estimation map to obtain the statistical quantity of the object to be counted in the image to be processed.

[0145] In some embodiments, the values at all positions in the density estimation map are between 0 and 1. When there is an object to be counted at a position in the image to be processed, the larger the value at this position in the density estimation map, the closer it is to 1. Among them, if there is no object to be counted in the image to be processed, the values at all positions in the corresponding density estimation map are 0; when there is an object to be counted at (X, Y) in the image to be processed, this object to be counted can affect the value of the affected area corresponding to (X, Y) in the density estimation map. For example, it can affect the values of all positions within the affected area with (X, Y) as the center and a radius of r, and the value at (X, Y) is the largest, and the values of other positions in the affected area become smaller as the distance from the center increases. It should be noted that in this affected area, the sum of the values at all positions is 1, that is, an object to be counted in the image to be processed can increase the sum of the values at all positions in the density estimation map by 1. Thus, the statistical quantity of the object to be counted in the image to be processed can be obtained by summing the density estimation map.

[0146] S5022, Obtain the local peak points in the density estimation map, and screen the local peak points through the non-maximum suppression algorithm to obtain the positions of the objects to be counted; the pixel values of the local peak points are greater than the pixel values of adjacent pixel points.

[0147] In some embodiments, since the larger the value at each position in the density estimation map, the higher the probability that there is an object to be counted at this position in the image to be processed, therefore, the preliminary positions of the objects to be counted can be determined by obtaining the local peak points in the density estimation map. Considering the case of errors, that is, for the same object to be counted in the image to be processed, there are at least two corresponding local peak points in the density estimation map. Therefore, it is necessary to use the non-maximum suppression algorithm to screen the obtained local peak points, and then obtain the positions of the objects to be counted.

[0148] Through the above embodiments, not only can the statistical quantity of the objects to be counted in the image to be processed be obtained, but also the positions of each object to be counted in the image to be processed can be obtained, expanding the application scope of the present disclosure.

[0149] Figure 6 It is an optional flowchart of the training method of the object statistical model provided by the embodiments of the present disclosure, and this method can be executed by the processor of a computer device. It will be described in combination with Figure 6 the steps shown.

[0150] Step S601: Obtain an image feature map corresponding to the sample image and at least one example feature vector corresponding to the object to be counted based on the feature acquisition network; the sample image includes a plurality of the objects to be counted.

[0151] In some embodiments, the feature acquisition network includes a feature extraction network and a feature processing network; the obtaining of the image feature map corresponding to the sample image and at least one example feature vector corresponding to the object to be counted based on the feature acquisition network includes: obtaining the sample image and at least one detection box corresponding to the sample image; the detection box is used to determine the area of the object to be counted in the sample image; extracting the image feature map corresponding to the sample image through the feature extraction network; based on the feature processing network, determining an example feature vector corresponding to each detection box through at least one detection box and the image feature map.

[0152] In some embodiments, the feature extraction network includes a feature fusion layer and a plurality of cascaded first convolutional layers; the extracting of the image feature map corresponding to the sample image through the feature extraction network includes: performing feature extraction processing on the sample image through the plurality of cascaded first convolutional layers to obtain an intermediate feature map output by each first convolutional layer; the sizes of the intermediate feature maps output by different first convolutional layers are different; fusing the intermediate feature maps output by each first convolutional layer through the feature fusion layer to obtain the image feature map.

[0153] In some embodiments, the feature processing network includes a feature truncation layer and a pooling layer; the determining of an example feature vector corresponding to each detection box through at least one detection box and the image feature map based on the feature processing network includes: obtaining the size ratio between the image feature map and the sample image through the feature truncation layer, performing scale transformation on each detection box based on the size ratio to obtain at least one feature truncation box; determining an example feature map corresponding to each detection box based on the at least one feature truncation box and the image feature map; converting the example feature map corresponding to each detection box through the pooling layer to obtain an example feature vector corresponding to each detection box.

[0154] Among them, global maximum pooling can be performed on the example feature map corresponding to each detection box to obtain an example feature vector corresponding to each detection box.

[0155] In some embodiments, the feature acquisition network includes a feature extraction network, a pooling layer, a feature fusion layer, and a plurality of cascaded first convolutional layers; obtaining the image feature map corresponding to the sample image and at least one example feature vector corresponding to the object to be counted based on the feature acquisition network includes: obtaining the sample image and at least one example image; the example image includes the object to be counted, and the example image is different from the sample image; performing feature extraction processing on the sample image through the feature extraction layer to obtain the image feature map corresponding to the sample image; performing feature extraction processing on at least one example image through the feature extraction layer to obtain at least one example feature map; converting the at least one example feature map into at least one example feature vector through the pooling layer.

[0156] Step S602, encoding the at least one example feature map through the encoding and decoding network to obtain a first feature sequence; decoding the first feature sequence and the image feature map to obtain the density matrix.

[0157] In some embodiments, the encoding and decoding network includes an encoding network and a decoding network; encoding the at least one example feature map through the encoding and decoding network to obtain a first feature sequence; decoding the first feature sequence and the image feature map to obtain the density matrix includes: encoding the at least one example feature vector through the encoding network to obtain a first feature sequence; decoding the first feature sequence and the image feature map through the decoding network to obtain the density matrix.

[0158] In some embodiments, the decoding network includes a first decoding layer and a second decoding layer; decoding the first feature sequence and the image feature map through the decoding network to obtain the density matrix includes: performing first decoding processing on the image feature map through the first decoding layer to obtain a second feature sequence; performing second decoding processing on the first feature sequence and the second feature sequence through the second decoding layer to obtain the density matrix.

[0159] In some embodiments, the second decoding layer includes a plurality of cascaded sub-decoding layers; the second decoding process of the first feature sequence and the second feature sequence through the second decoding layer to obtain the density matrix includes: taking the second feature sequence as the input feature corresponding to the first sub-decoding layer among the plurality of sub-decoding layers; for each sub-decoding layer, inputting the first feature sequence and the input feature corresponding to the sub-decoding layer into the sub-decoding layer to obtain a third feature sequence output by the sub-decoding layer, and taking the third feature sequence output by the sub-decoding layer as the input feature corresponding to the next sub-decoding layer; wherein, the density matrix is determined based on the third feature sequence output by the last sub-decoding layer among the plurality of sub-decoding layers.

[0160] In some embodiments, the encoding layer includes a plurality of cascaded sub-encoding layers; the encoding process of the at least one example feature vector through the encoding network to obtain a first feature sequence includes: taking the at least one example feature vector as the input feature corresponding to the first sub-encoding layer among the plurality of sub-encoding layers; for each sub-encoding layer, inputting the input feature corresponding to the sub-encoding layer into the sub-encoding layer to obtain a fourth feature sequence output by the sub-encoding layer, and taking the fourth feature sequence output by the sub-encoding layer as the input feature corresponding to the next sub-encoding layer; wherein, the first feature sequence is the fourth feature sequence output by the last sub-encoding layer among the plurality of sub-encoding layers.

[0161] In this embodiment, the encoding and decoding network is an encoding and decoding network based on the attention mechanism. In some embodiments, the encoding and decoding network can be but is not limited to a Transformer network. Correspondingly, the encoding network in the encoding and decoding network is the encoder of the Transformer network, and the decoding network in the encoding and decoding network is the decoder of the Transformer network. The above step S602 can be implemented in the following manner: encoding the at least one example feature vector based on the encoder of the Transformer network to obtain a first feature sequence; decoding the first feature sequence and the image feature map based on the decoder of the Transformer network to obtain a density matrix.

[0162] Step S603, converting the density matrix into a predicted density map through the density regression network.

[0163] In some embodiments, the density regression network is used to convert the density matrix into a predicted density map. The density regression network may include a plurality of 1x1 convolutional layers and a resampling layer. Through the plurality of 1x1 convolutional layers, the density matrix can be converted into an intermediate density map with 1 channel, and then the intermediate density map is resampled based on the resampling layer to obtain a density estimation map. The size of the density estimation map is the same as the size of the image to be processed.

[0164] Step S604: Based on the predicted density map and the true density map corresponding to the sample image, adjust the parameters of the object statistical model to be trained to obtain a trained object statistical model.

[0165] The object statistical model obtained by the object statistical model training method provided by the above embodiments can encode the at least one example feature vector through a self-attention mechanism to obtain a first feature sequence, and can focus on the common features of the object to be counted under different examples on the basis of each example feature vector; at the same time, since the density matrix is obtained by decoding the first feature sequence and the image feature map, the obtained density matrix can focus on the common features of the object to be counted under different examples on the basis of the image feature map, so that the accuracy of the statistical result can be improved in the process of generating the statistical result for the object to be counted based on the density matrix.

[0166] Figure 7 is an optional flowchart of the model training method provided by the embodiments of the present disclosure, and this method can be executed by the processor of a computer device. Based on Figure 1 , Figure 1 It may further include S701 to S703, which will be described in conjunction with Figure 7 the steps shown.

[0167] Step S701: Construct an initial matrix; the size of the initial matrix is the same as the size of the sample image.

[0168] In some embodiments, the element value at each matrix position in the initial matrix can be set to a preset first value.

[0169] Step S702: Obtain annotation information for the sample image; the annotation information is used to represent the relative coordinates of each object to be counted in the sample image.

[0170] In some embodiments, for each object to be counted, the relative coordinates of the feature points of the object to be counted in the sample image are obtained. In some embodiments, the feature points of the object to be counted may be the center point of the detection frame of the object to be counted; in other embodiments, the feature points of the object to be counted may also be the feature part points of the object to be counted. Exemplarily, when the object to be counted is a flower, the feature part point may be the point where the flower center is located; when the object to be counted is a bird, the feature part point may be the point where the bird's head is located.

[0171] Step S703: Based on the relative coordinates of each object to be counted in the sample image, update the element value of the matrix position corresponding to the initial matrix to obtain the true density map.

[0172] In some embodiments, for each object to be counted, based on the relative coordinates of the object to be counted in the sample image, determine the matrix position corresponding to the object to be counted in the initial matrix; update the first value at the matrix position in the initial matrix to a preset second value.

[0173] Wherein, the first value can be set to 0, and the second value can be set to 1. At this time, when there are T objects to be counted in the sample image, the sum of the element values of each matrix position in the updated initial matrix is also T.

[0174] In some embodiments, after updating the element value of the matrix position corresponding to the initial matrix based on the relative coordinates of each object to be counted in the sample image, the updated initial matrix can also be subjected to convolution processing using a preset Gaussian kernel to obtain the true density map. At this time, when there are T objects to be counted in the sample image, the sum of the element values of each matrix position in the updated initial matrix is still T.

[0175] The following describes the application of the object counting method provided by the embodiments of the present disclosure in an actual scenario. The object counting method may be implemented by multiple modules. In some embodiments, the multiple modules may include a data acquisition module, a network structure module, a training module, and an inference module.

[0176] In some embodiments, the data acquisition module is used to acquire the image to be processed. The image to be processed may be all data existing in the form of an image, including photos, remote sensing images, electron microscope images, etc. Among them, the image to be processed may include multiple objects to be counted, and the objects to be counted may include but are not limited to object objects, human objects, and other biological objects, etc. For example, the input data is an image data A, and the image data A may contain multiple (1 to 5000) same objects (such as people, cars, cells, etc.). There may be differences between different individuals of the same object.

[0177] Among them, during the implementation of the data acquisition module, it can be divided into two scenarios: "training" and "inference".

[0178] In the "training" scenario, on the one hand, a small number (<= 5) of rectangular bounding box annotations are made for the objects to be counted on the real image manually to obtain the first sample image. The first sample image carries at least one bounding box information, and each bounding box information is used to determine the position of the object to be counted in the first sample image.

[0179] On the other hand, it is also necessary to obtain the corresponding real density image of the first sample image as the second sample image. The method for obtaining the real density image includes: generating a first matrix based on the image size of the first sample image, the size of the first matrix is the same as the image size of the first sample image, and the values in the first matrix are all 0; obtaining the position information of each object to be counted in the first sample image, updating the values in the first matrix corresponding to the position information of each object to be counted to 1 to obtain a second matrix; performing a blur process on the second matrix to obtain the real density image (second sample image).

[0180] It should be noted that the position information of the object to be counted is the coordinate information of the feature points of the object to be counted. In some embodiments, the feature points of the object to be counted can be the center point of the detection box of the object to be counted; in other embodiments, the feature points of the object to be counted can also be the feature part points of the object to be counted. Exemplarily, when the object to be counted is a flower, the feature part point can be the point where the flower center is located; when the object to be counted is a bird, the feature part point can be the point where the bird's head is located.

[0181] In some embodiments, the above-mentioned blur process on the second matrix to obtain the real density image can be achieved by the following method: performing a convolution process on the second matrix using a preset Gaussian convolution kernel to obtain the real density image.

[0182] In some implementation scenarios, please refer to Figure 8a , which shows a real image. In this real image 81, it includes a plurality of objects to be counted, and exemplarily, it can include 811, 812, 813, etc. Please refer to Figure 8b , which shows a schematic diagram of a rectangular bounding box. In the real image 81, corresponding rectangular bounding boxes can be added to at least one object to be counted. Exemplarily, Figure 8b 3 corresponding rectangular bounding boxes are added to the objects to be counted. Among them, a corresponding rectangular bounding box 821 is added to the object to be counted 811, a corresponding rectangular bounding box 822 is added to the object to be counted 812, and a corresponding rectangular bounding box 823 is added to the object to be counted 813. Please refer to Figure 8c, which shows a schematic diagram of a true density map. It can be seen that each object to be counted in the true image 81 corresponds to a density point. For example, the object to be counted 811 corresponds to the density point 831 in the true density map, the object to be counted 812 corresponds to the density point 832 in the true density map, and the object to be counted 813 corresponds to the density point 833 in the true density map.

[0183] In the "inference" scenario, it is necessary to manually annotate a small number (<= 5) of rectangular bounding boxes for the objects to be counted on the image to be processed.

[0184] In some embodiments, please refer to Figure 8d the schematic diagram of the network structure module shown, which may include a feature extraction module 841, a feature processing module 842, a feature reconstruction module 843, a density regression module 844, and a post-processing module 845.

[0185] In some embodiments, the feature extraction module 841 is used to extract the image feature map corresponding to the image to be processed. Among them, the feature extraction network can be a pre-trained convolutional neural network, including but not limited to Resnet, EfficientNet, VGG, etc. Among them, the convolutional neural network can include a plurality of sequentially connected convolutional layers. In the inference process after inputting the image to be processed into the convolutional neural network, each convolutional layer can output a sub-feature map of a scale, obtaining multiple scales of sub-feature maps. Resample some of the sub-feature maps of multiple scales to a unified size, and splice some of the sub-feature maps of multiple scales based on the channel direction to obtain the finally output multi-scale feature map, that is, the image feature map corresponding to the image to be processed. This multi-scale feature Figure 1 is input into the feature processing module 842 on the one hand and the feature reconstruction module 843 on the other hand.

[0186] In some embodiments, the feature processing module 842 is used to determine the example feature map corresponding to each bounding box. Among them, each bounding box is used to determine the object to be counted in the image to be processed. Correspondingly, the example feature map corresponding to the bounding box carries the feature information of the object to be counted in the bounding box. Therefore, based on the bounding box information corresponding to each bounding box, the example feature map corresponding to each bounding box can be cropped in the image feature map. After that, the example feature map corresponding to each bounding box can be used as the query feature vector Q and input into the feature reconstruction module 843.

[0187] Exemplarily, after receiving the bounding box information of multiple bounding boxes, the feature processing module 842 crops the sub-feature map of the object to be counted in the bounding box in the image feature map based on the bounding box information (the position of the bounding box). For each sub-feature map, global maximum pooling is performed on the sub-feature map to obtain the example feature map corresponding to the sub-feature map.

[0188] Taking the bounding box with 3 annotations as an example, 3 sub-feature maps can be obtained at the corresponding positions of the image feature map corresponding to the image to be processed. For the 3 sub-feature maps, global maximum pooling is performed respectively to obtain 3 example feature maps.

[0189] In some embodiments, the feature reconstruction module is used to generate a density matrix based on the image feature map corresponding to the image to be processed and the example feature map corresponding to each bounding box. Among them, the feature reconstruction module 843 may include a Transformer network. Correspondingly, the Transformer network includes two parts: a Transformer encoder and a Transformer decoder. After taking the image feature map obtained by the feature extraction module 841 as the input of the Transformer decoder and taking the example feature map corresponding to each bounding box obtained by the feature processing module 842 as the input of the Transformer encoder, the image feature map can be input into the density regression module 844. The feature reconstruction module 843 can output a density matrix. The size of the density matrix is the same as the size of the image feature map.

[0190] In some embodiments, the density regression module 844 is used to convert the density matrix into a density estimation map. Among them, the density regression module 844 includes a plurality of 1x1 convolutional layers and a resampling layer. Inputting the density matrix into the plurality of 1x1 convolutional layers can obtain an intermediate density map with 1 channel, and then resampling the intermediate density map based on the resampling layer to obtain a density estimation map. The size of the density estimation map is the same as the size of the image to be processed.

[0191] In some embodiments, the post-processing module 845 is used to generate the statistical result corresponding to the image to be processed based on the density estimation map. Among them, the pixel values of each pixel point in the density estimation map can be accumulated, and the sum obtained is the statistical quantity of the object to be statistically counted in the image to be processed; based on the pixel values of each pixel point in the density estimation map, a plurality of local peak points (i.e., points with pixel values greater than the surrounding 8 pixels) and the position coordinates of each local peak point in the density estimation map are determined, and non-maximum suppression (NMS) is performed on the plurality of local peak points, and the coordinates of each remaining local peak point are used as the position of each object to be statistically counted.

[0192] In some embodiments, the input of the training module is a sample image, the corresponding indication of the object bounding box, and the ground-truth density map. Taking the network of the above-mentioned module as the object, an object statistical model is finally trained. During training, the image and the bounding box are input into the network, and the mean square error loss function is calculated between the density estimation map B' output by the network and the ground-truth density map B. For each image, the obtained loss value Loss is used for the gradient descent optimization method to optimize the parameter weights in the network. By iteratively training multiple times according to this method until a predetermined number of iterations is reached.

[0193] In some embodiments, the inference module is the module used in actual applications. It uses the trained object statistical model obtained by the training module to estimate the count of a specified object in any image. The input image is any to-be-processed image containing multiple objects of the same class to be counted, and the object bounding box indicating the object to be counted. The output is the estimated value of the number of objects to be counted and the position in the image.

[0194] Based on the foregoing embodiments, the embodiments of the present disclosure provide an object counting device and a training device for an object statistical model. The device includes each unit included, and each module included in each unit, and can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits; during implementation, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0195] Figure 9 FIG. is a schematic structural diagram of an object counting device provided by an embodiment of the present disclosure, as Figure 9 shown, the object counting device 900 includes: a first acquisition module 901, an encoding module 902, a decoding module 903, and a generation module 904, wherein:

[0196] The first acquisition module 901 is configured to acquire an image feature map corresponding to the to-be-processed image and at least one example feature vector corresponding to the object to be counted; the to-be-processed image includes multiple objects to be counted;

[0197] The encoding module 902 is configured to perform encoding processing on the at least one example feature vector to obtain a first feature sequence;

[0198] The decoding module 903 is configured to perform decoding processing on the first feature sequence and the image feature map to obtain a density matrix;

[0199] A generation module 904, configured to generate a statistical result for the object to be counted in the image to be processed based on the density matrix.

[0200] In some embodiments, the decoding module 903 is further configured to: perform a first decoding process on the image feature map to obtain a second feature sequence; and perform a second decoding process on the first feature sequence and the second feature sequence to obtain the density matrix.

[0201] In some embodiments, the second decoding process includes multiple iterative sub-decoding processes; the decoding module 903 is further configured to: use the second feature sequence as the input feature corresponding to the first sub-decoding process in the multiple sub-decoding processes; for each sub-decoding process, use the first feature sequence and the input feature corresponding to the sub-decoding process as the input of the sub-decoding process to obtain a third feature sequence output by the sub-decoding process, and use the third feature sequence output by the sub-decoding process as the input feature corresponding to the next sub-decoding process; wherein, the density matrix is determined based on the third feature sequence output by the last sub-decoding process in the multiple sub-decoding processes.

[0202] In some embodiments, the encoding process includes multiple iterative sub-encoding processes; the encoding module 902 is further configured to: use the at least one example feature vector as the input feature corresponding to the first sub-encoding process in the multiple sub-encoding processes; for each sub-encoding process, use the input feature corresponding to the sub-encoding process as the input of the sub-encoding process to obtain a fourth feature sequence output by the sub-encoding process, and use the fourth feature sequence output by the sub-encoding process as the input feature corresponding to the next sub-encoding process; wherein, the first feature sequence is the fourth feature sequence output by the last sub-encoding process in the multiple sub-encoding processes.

[0203] In some embodiments, the first acquisition module 901 is further configured to: acquire the image to be processed and at least one detection box corresponding to the image to be processed; the detection box is used to determine the area of the object to be counted in the image to be processed; extract the image feature map corresponding to the image to be processed; and determine an example feature vector corresponding to each detection box based on at least one detection box and the image feature map.

[0204] In some embodiments, the first acquisition module 901 is further configured to: perform feature extraction processing on the image to be processed at multiple scales to obtain an intermediate feature map corresponding to each scale; the sizes of the intermediate feature maps at different scales are different; and fuse the intermediate feature maps corresponding to each scale to obtain the image feature map.

[0205] In some embodiments, the first acquisition module 901 is further configured to: obtain the size ratio between the image feature map and the image to be processed; perform scale transformation on each of the detection frames based on the size ratio to obtain at least one feature extraction frame; determine an example feature map corresponding to each of the detection frames based on the at least one feature extraction frame and the image feature map; and determine an example feature vector corresponding to each of the detection frames based on the example feature map corresponding to each of the detection frames.

[0206] In some embodiments, the first acquisition module 901 is further configured to: obtain the image to be processed and at least one example image; the example image includes the object to be counted, and the example image is different from the image to be processed; extract the image feature map corresponding to the image to be processed; and extract an example feature vector corresponding to each of the example images.

[0207] In some embodiments, the generation module 904 is further configured to: generate a density estimation map based on the density matrix; and generate a statistical result for the object to be counted in the image to be processed based on the density estimation map.

[0208] In some embodiments, the generation module 904 is further configured to: perform iterative convolution processing on the density matrix; and resample the density matrix after the iterative convolution processing to obtain a density estimation map having the same size as the image to be processed.

[0209] In some embodiments, the generation module 904 is further configured to: sum the density estimation map to obtain the statistical quantity of the object to be counted in the image to be processed; and / or, obtain local peak points in the density estimation map, and screen the local peak points through a non-maximum suppression algorithm to obtain the position of the object to be counted; the pixel value of the local peak point is greater than the pixel values of adjacent pixel points.

[0210] Figure 10 FIG. is a schematic structural diagram of a training device for an object statistical model provided by an embodiment of the present disclosure. The object statistical model includes a feature acquisition network, an encoding and decoding network, and a density regression network; as Figure 10 shown, the training device 1000 of the object statistical model includes: a second acquisition module 1001, an encoding and decoding module 1002, a regression module 1003, and an adjustment module 1004, where:

[0211] The second acquisition module 1001 is configured to obtain an image feature map corresponding to a sample image and at least one example feature vector corresponding to an object to be counted based on the feature acquisition network; the sample image includes a plurality of the objects to be counted;

[0212] The encoding and decoding module 1002 is configured to perform encoding processing on the at least one example feature map through the encoding and decoding network to obtain a first feature sequence; and perform decoding processing on the first feature sequence and the image feature map to obtain the density matrix.

[0213] The regression module 1003 is configured to convert the density matrix into a predicted density map through the density regression network.

[0214] The adjustment module 1004 is configured to adjust the parameters of the object statistical model to be trained based on the predicted density map and the ground truth density map corresponding to the sample image, so as to obtain a trained object statistical model.

[0215] In some embodiments, the second acquisition module 1001 is further configured to: construct an initial matrix; the size of the initial matrix is the same as the size of the sample image; obtain annotation information for the sample image; the annotation information is used to represent the relative coordinates of each object to be counted in the sample image; and update the element values at the matrix positions corresponding to the initial matrix based on the relative coordinates of each object to be counted in the sample image to obtain the ground truth density map.

[0216] In some embodiments, the encoding and decoding network includes an encoding network and a decoding network; the encoding and decoding module 1002 is further configured to: perform encoding processing on the at least one example feature vector through the encoding network to obtain a first feature sequence; and perform decoding processing on the first feature sequence and the image feature map through the decoding network to obtain the density matrix.

[0217] In some embodiments, the decoding network includes a first decoding layer and a second decoding layer; the encoding and decoding module 1002 is further configured to: perform first decoding processing on the image feature map through the first decoding layer to obtain a second feature sequence; and perform second decoding processing on the first feature sequence and the second feature sequence through the second decoding layer to obtain the density matrix.

[0218] In some embodiments, the second decoding layer includes a plurality of cascaded sub-decoding layers; the encoding and decoding module 1002 is further configured to: use the second feature sequence as the input feature corresponding to the first sub-decoding layer among the plurality of sub-decoding layers; for each sub-decoding layer, input the first feature sequence and the input feature corresponding to the sub-decoding layer into the sub-decoding layer to obtain a third feature sequence output by the sub-decoding layer, and use the third feature sequence output by the sub-decoding layer as the input feature corresponding to the next sub-decoding layer; wherein, the density matrix is determined based on the third feature sequence output by the last sub-decoding layer among the plurality of sub-decoding layers.

[0219] In some embodiments, the encoding layer includes a plurality of cascaded sub-encoding layers; the encoding and decoding module 1002 is further configured to: use the at least one example feature vector as the input feature corresponding to the first sub-encoding layer among the plurality of sub-encoding layers; for each of the sub-encoding layers, input the input feature corresponding to the sub-encoding layer into the sub-encoding layer to obtain a fourth feature sequence output by the sub-encoding layer, and use the fourth feature sequence output by the sub-encoding layer as the input feature corresponding to the next sub-encoding layer; wherein, the first feature sequence is the fourth feature sequence output by the last sub-encoding layer among the plurality of sub-encoding layers.

[0220] In some embodiments, the feature acquisition network includes a feature extraction network and a feature processing network; the encoding and decoding module 1002 is further configured to: acquire the sample image and at least one detection box corresponding to the sample image; the detection box is used to determine the region of the object to be counted in the sample image; extract an image feature map corresponding to the sample image through the feature extraction network; based on the feature processing network, determine an example feature vector corresponding to each detection box through at least one of the detection boxes and the image feature map.

[0221] In some embodiments, the feature extraction network includes a feature fusion layer and a plurality of cascaded first convolutional layers; the second acquisition module 1001 is further configured to: perform feature extraction processing on the sample image through the plurality of cascaded first convolutional layers to obtain an intermediate feature map output by each of the first convolutional layers; the sizes of the intermediate feature maps output by different first convolutional layers are different; fuse the intermediate feature maps output by each of the first convolutional layers through the feature fusion layer to obtain the image feature map.

[0222] In some embodiments, the feature processing network includes a feature truncation layer and a pooling layer; the second acquisition module 1001 is further configured to: obtain the size ratio between the image feature map and the sample image through the feature truncation layer, perform scale transformation on each detection box based on the size ratio to obtain at least one feature truncation box; determine an example feature map corresponding to each detection box based on the at least one feature truncation box and the image feature map; perform conversion on the example feature map corresponding to each detection box through the pooling layer to obtain an example feature vector corresponding to each detection box.

[0223] In some embodiments, the feature acquisition network includes a feature extraction network, a pooling layer, a feature fusion layer, and a plurality of cascaded first convolutional layers; the second acquisition module 1001 is further configured to: acquire the sample image and at least one example image; the example image includes the object to be counted, and the example image is different from the sample image; perform feature extraction processing on the sample image through the feature extraction layer to obtain an image feature map corresponding to the sample image; perform feature extraction processing on at least one example image through the feature extraction layer to obtain at least one example feature map; convert the at least one example feature map into at least one example feature vector through the pooling layer.

[0224] The description of the above device embodiments is similar to the description of the above method embodiments and has similar beneficial effects to the method embodiments. In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. For the technical details not disclosed in the device embodiments of the present disclosure, please refer to the description of the method embodiments of the present disclosure for understanding.

[0225] It should be noted that in the embodiments of the present disclosure, if the above object counting method is implemented in the form of software function modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present disclosure essentially or the part that contributes to the related technology can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present disclosure. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes. In this way, the embodiments of the present disclosure are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.

[0226] The embodiments of the present disclosure provide a computer device, including a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the program, it implements some or all of the steps in the above method.

[0227] The embodiments of the present disclosure provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements some or all of the steps in the above method. The computer-readable storage medium can be transient or non-transient.

[0228] An embodiment of the present disclosure provides a computer program, including computer-readable code. When the computer-readable code runs on a computer device, a processor in the computer device executes to implement some or all of the steps in the above method.

[0229] An embodiment of the present disclosure provides a computer program product. The computer program product includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, some or all of the steps in the above method are implemented. The computer program product can be specifically implemented in a manner of hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium. In other embodiments, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.

[0230] It should be noted here that: the descriptions of the above embodiments tend to emphasize the differences between the embodiments, and their similarities can be referred to each other. The descriptions of the above device, storage medium, computer program, and computer program product embodiments are similar to the descriptions of the above method embodiments and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of the present disclosure, please refer to the descriptions of the method embodiments of the present disclosure for understanding.

[0231] Figure 11 The following is a schematic diagram of the hardware entity of a computer device provided by an embodiment of the present disclosure. As Figure 11 shown, the hardware entity of the computer device 1100 includes: a processor 1101 and a memory 1102. Among them, the memory 1102 stores a computer program that can run on the processor 1101. When the processor 1101 executes the program, it implements the steps in the method of any of the above embodiments.

[0232] The memory 1102 stores a computer program that can run on the processor. The memory 1102 is configured to store instructions and applications executable by the processor 1101, and can also cache data to be processed or already processed by the processor 1101 and each module in the computer device 1100 (for example, image data, audio data, voice communication data, and video communication data), and can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM).

[0233] When the processor 1101 executes the program, it implements the steps of the object statistics method of any of the above items. The processor 1101 generally controls the overall operation of the computer device 1100.

[0234] Embodiments of the present disclosure provide a computer storage medium storing one or more programs that can be executed by one or more processors to implement the steps of the object statistics method or the training method of the object statistics model in any of the foregoing embodiments.

[0235] It should be noted here that the descriptions of the above storage medium and device embodiments are similar to those of the foregoing method embodiments and have similar beneficial effects to the method embodiments. For the technical details not disclosed in the storage medium and device embodiments of the present disclosure, please refer to the descriptions of the method embodiments of the present disclosure for understanding.

[0236] The foregoing processor may be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor. It can be understood that other electronic devices capable of implementing the functions of the foregoing processor are also possible, and the embodiments of the present disclosure do not make specific limitations.

[0237] The foregoing computer storage medium / memory may be a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM), etc.; it may also be various terminals including one or any combination of the foregoing memories, such as a mobile phone, a computer, a tablet device, a personal digital assistant, etc.

[0238] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present disclosure. Therefore, the "in one embodiment" or "in an embodiment" that appears throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present disclosure, the magnitude of the sequence numbers of the above steps / processes does not mean the order of execution. The order of execution of each step / process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure. The sequence numbers of the embodiments of the present disclosure above are only for description and do not represent the advantages or disadvantages of the embodiments.

[0239] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0240] In several embodiments provided by the present disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling or communication connection between the components shown or discussed with each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0241] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units; they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0242] In addition, each functional unit in the embodiments of the present disclosure may be entirely integrated into one processing unit, or each unit may be individually regarded as one unit, or two or more units may be integrated into one unit; the above integrated unit may be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0243] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: removable storage devices, read-only memory (ROM), magnetic disks, or optical disks and other various media that can store program codes.

[0244] Alternatively, if the above integrated unit of the present disclosure is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the related technology can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the embodiments of the present disclosure. And the foregoing storage medium includes: removable storage devices, ROM, magnetic disks, or optical disks and other various media that can store program codes.

[0245] As described above, only the embodiments of the present disclosure are provided, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of changes or substitutions, which should all be covered by the protection scope of the present disclosure.

Claims

1. An object counting method, characterized in that, The method includes: Obtaining an image feature map corresponding to the image to be processed and at least one example feature vector corresponding to the object to be counted; the image to be processed includes a plurality of the objects to be counted; the at least one example feature vector is used to characterize the feature information of the object to be counted in different example images; the example image corresponding to the object to be counted is an image including the object to be counted; Performing encoding processing on the at least one example feature vector to obtain a first feature sequence; the first feature sequence is used to characterize the common features of the object to be counted in different example images; Performing decoding processing on the first feature sequence and the image feature map to obtain a density matrix; Generating a statistical result for the object to be counted in the image to be processed based on the density matrix.

2. The method according to claim 1, wherein The performing decoding processing on the first feature sequence and the image feature map to obtain a density matrix includes: Performing first decoding processing on the image feature map to obtain a second feature sequence; Performing second decoding processing on the first feature sequence and the second feature sequence to obtain the density matrix.

3. The method according to claim 2, wherein The second decoding processing includes multiple iterative sub-decoding processes; the performing second decoding processing on the first feature sequence and the second feature sequence to obtain the density matrix includes: Using the second feature sequence as the input feature corresponding to the first sub-decoding process in the multiple sub-decoding processes; For each of the sub-decoding processes, using the first feature sequence and the input feature corresponding to the sub-decoding process as the input of the sub-decoding process to obtain a third feature sequence output by the sub-decoding process, and using the third feature sequence output by the sub-decoding process as the input feature corresponding to the next sub-decoding process; Wherein, the density matrix is determined based on the third feature sequence output by the last sub-decoding process in the multiple sub-decoding processes.

4. The method according to any one of claims 1 to 3, characterized in that, The obtaining an image feature map corresponding to the image to be processed and at least one example feature vector corresponding to the object to be counted further includes: Obtaining the image to be processed and at least one detection box corresponding to the image to be processed; the detection box is used to determine the region of the object to be counted in the image to be processed; Extracting the image feature map corresponding to the image to be processed; Determining an example feature vector corresponding to each detection box based on at least one of the detection boxes and the image feature map.

5. The method according to claim 4, characterized in that, The extracting the image feature map corresponding to the image to be processed includes: Performing feature extraction processing on the image to be processed at multiple scales to obtain an intermediate feature map corresponding to each scale; the sizes of the intermediate feature maps at different scales are different; Fusing the intermediate feature maps corresponding to each scale to obtain the image feature map.

6. The method according to claim 4, characterized in that The determining an example feature vector corresponding to each detection box based on at least one of the detection boxes and the image feature map includes: Obtaining the size ratio between the image feature map and the image to be processed; Performing scale transformation on each detection box based on the size ratio to obtain at least one feature extraction box; Determine an example feature map corresponding to each of the detection boxes based on the at least one feature intercepting box and the image feature map; Determine an example feature vector corresponding to each of the detection boxes based on the example feature map corresponding to each of the detection boxes.

7. The method according to any one of claims 1 to 3, characterized in that, The obtaining of the image feature map corresponding to the image to be processed and at least one example feature vector corresponding to the object to be counted includes: Obtain the image to be processed and at least one example image; the example image is different from the image to be processed; Extract the image feature map corresponding to the image to be processed; Extract the example feature vector corresponding to each of the example images.

8. The method according to any one of claims 1 to 3, characterized in that The generating of the statistical result for the object to be counted in the image to be processed based on the density matrix includes: Generate a density estimation map based on the density matrix; Generate the statistical result for the object to be counted in the image to be processed based on the density estimation map.

9. The method according to claim 8, wherein The generating of the density estimation map based on the density matrix includes: performing iterative convolution processing on the density matrix; resampling the density matrix after the iterative convolution processing to obtain a density estimation map having the same size as the image to be processed; The generating of the statistical result for the object to be counted in the image to be processed based on the density estimation map includes: summing the density estimation map to obtain the statistical quantity of the object to be counted in the image to be processed; and / or, obtaining local peak points in the density estimation map, and screening the local peak points by using a non-maximum suppression algorithm to obtain the position of the object to be counted; the pixel value of the local peak point is greater than the pixel values of adjacent pixel points.

10. A training method for an object statistical model, characterized in that, The statistical model includes a feature acquisition network, an encoding and decoding network, and a density regression network, and the method includes: Obtain the image feature map corresponding to the sample image and at least one example feature vector corresponding to the object to be counted based on the feature acquisition network; the sample image includes a plurality of the objects to be counted; the at least one example feature vector is used to characterize the feature information of the object to be counted in different example images; the example image corresponding to the object to be counted is an image including the object to be counted; Perform encoding processing on the at least one example feature map through the encoding and decoding network to obtain a first feature sequence; the first feature sequence is used to characterize the common features of the object to be counted in different example images; perform decoding processing on the first feature sequence and the image feature map to obtain a density matrix; Convert the density matrix into a predicted density map through the density regression network; Adjust the parameters of the object statistical model to be trained based on the predicted density map and the true density map corresponding to the sample image to obtain a trained object statistical model.

11. The method according to claim 10, wherein The method further includes: Construct an initial matrix; the size of the initial matrix is the same as the size of the sample image; Obtain the annotation information for the sample image; the annotation information is used to characterize the relative coordinates of each of the objects to be counted in the sample image; Updating the element value of the matrix position corresponding to the initial matrix based on the relative coordinates of each of the objects to be counted in the sample image to obtain the true density map.

12. An object counting device, characterized in that, Including: A first acquisition module, configured to acquire an image feature map corresponding to an image to be processed and at least one example feature vector corresponding to an object to be counted; the image to be processed includes a plurality of the objects to be counted; the at least one example feature vector is used to characterize the feature information of the object to be counted in different example images; the example image corresponding to the object to be counted is an image including the object to be counted; An encoding module, configured to perform encoding processing on the at least one example feature vector to obtain a first feature sequence; the first feature sequence is used to characterize the common features of the object to be counted in different example images; A decoding module, configured to perform decoding processing on the first feature sequence and the image feature map to obtain a density matrix; A generation module, configured to generate a statistical result for the object to be counted in the image to be processed based on the density matrix.

13. A computer device, comprising a memory and a processor, the memory storing a computer program that can run on the processor, characterized in that, When the processor executes the program, it implements the steps in the method according to any one of claims 1 to 9, or implements the steps in the method according to any one of claims 10 to 11 when executing the program.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps in the method according to any one of claims 1 to 9, or implements the steps in the method according to any one of claims 10 to 11.

Citation Information

Patent Citations

  • Image processing method and device, model training method and device, equipment and storage medium

    CN114612414A