Image processing method and device
By introducing dual loss function training focusing on pixel point density and area object number in the image processing model, the problem of low accuracy of group objects in the prior art is solved, and higher accuracy and effectiveness are achieved.
Patent Information
- Application Number
- CN202110113954.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-27
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2041-01-27
AI Technical Summary
In the prior art, the image processing model only focuses on the prediction of the population density of pixels during the training process, resulting in the accuracy of the final group object number and the training target is inconsistent with the usage target.
The first loss function is used to focus on the population density of pixel points, and the second loss function is used to focus on the number of population objects in the image area. The target image processing model is trained in combination with the two to improve the prediction accuracy of the model.
Through rich training information, the problem of inconsistency between training goals and usage goals is alleviated, and the accuracy of determining the number of group objects is improved and the processing effect is improved.
Smart Images

Figure CN114898282B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular to an image processing method and device. Background Art
[0002] With the development of artificial intelligence (AI) technology, using AI to analyze the number of groups corresponding to an image (i.e., the number of groups contained in an image) has played a significant role in areas such as public safety. Currently, this method involves processing an image using an image processing model to obtain a group density map corresponding to the image. The group density map then determines the number of groups corresponding to the image. The group density map indicates the group density corresponding to each pixel in the image.
[0003] In related art, image processing models are trained directly using a loss function determined based on the predicted population density corresponding to each sample pixel in a sample image. This approach focuses solely on whether the image processing model can accurately predict the population density corresponding to each sample pixel in the sample image, which provides relatively limited information.
[0004] The ultimate goal of the image processing model is to accurately predict a population density map that can accurately determine the number of population objects corresponding to an image. In a training method that only focuses on whether the image processing model can accurately predict the population density corresponding to each sample pixel point, it is easy for the training goal to be inconsistent with the ultimate goal. This will result in a low accuracy in the number of population objects corresponding to the image determined by the population density map obtained by the trained image processing model, and a poor processing effect when calling the trained image processing model to process the image. Summary of the Invention
[0005] The present application provides an image processing method and apparatus that can be used to improve the accuracy of determining the number of group objects corresponding to an image. The technical solution is as follows:
[0006] In one aspect, an embodiment of the present application provides an image processing method, the method comprising:
[0007] Obtaining a target image processing model and a target image to be processed, wherein the target image processing model is trained using a first loss function and a second loss function, wherein the first loss function is determined based on a predicted population density corresponding to each sample pixel point in the sample image, and the second loss function is determined based on a predicted number of population objects corresponding to each first image region in the sample image;
[0008] Calling the target image processing model to process the target image to obtain a target group density map corresponding to the target image;
[0009] The number of target group objects corresponding to the target image is determined based on the target group density map, where the number of target group objects corresponding to the target image is used to indicate the number of group objects contained in the target image.
[0010] In another aspect, an image processing apparatus is provided, the apparatus comprising:
[0011] an acquisition unit, configured to acquire a target image processing model and a target image to be processed, wherein the target image processing model is trained using a first loss function and a second loss function, wherein the first loss function is determined based on a predicted population density corresponding to each sample pixel point in the sample image, and the second loss function is determined based on a predicted number of population objects corresponding to each first image region in the sample image;
[0012] a calling unit, configured to call the target image processing model to process the target image, and obtain a target group density map corresponding to the target image;
[0013] A determining unit is configured to determine the number of target group objects corresponding to the target image based on the target group density map, where the number of target group objects corresponding to the target image is used to indicate the number of group objects contained in the target image.
[0014] In a possible implementation, the acquisition unit is further configured to acquire a sample image and a standard population density map corresponding to the sample image, wherein the standard population density map is configured to indicate the standard population density corresponding to each sample pixel point in the sample image.
[0015] The calling unit is further configured to call the initial image processing model to process the sample image to obtain a predicted population density map corresponding to the sample image, wherein the predicted population density map is used to indicate the predicted population density corresponding to each sample pixel point in the sample image;
[0016] The determining unit is further configured to determine a first loss function based on the predicted population density corresponding to each sample pixel point and the standard population density corresponding to each sample pixel point;
[0017] The device further comprises:
[0018] a dividing unit, configured to divide the sample image into a first reference number of first image regions;
[0019] The determining unit is further configured to determine the predicted number of group objects corresponding to the first reference number of first image regions and the standard number of group objects corresponding to the first reference number of first image regions;
[0020] The determining unit is further configured to determine a second loss function based on the predicted number of group objects corresponding to the first reference number of first image regions and the standard number of group objects corresponding to the first reference number of first image regions;
[0021] The device further comprises:
[0022] A training unit is used to train the initial image processing model using the first loss function and the second loss function to obtain the target image processing model.
[0023] In one possible implementation, the determination unit is further used to determine a first target area in the first image area of the first reference quantity based on the predicted number of group objects corresponding to the first image area of the first reference quantity and the standard number of group objects corresponding to the first image area of the first reference quantity; determine candidate pixel points in the first target area; and determine a second loss function based on the predicted group density corresponding to the candidate pixel points and the standard group density corresponding to the candidate pixel points.
[0024] In one possible implementation, the determination unit is further configured to, for any first image area among the first reference number of first image areas, in response to the predicted number of group objects corresponding to any first image area being greater than the standard number of group objects corresponding to any first image area, use the first type as the area type corresponding to any first image area; in response to the predicted number of group objects corresponding to any first image area being less than the standard number of group objects corresponding to any first image area, use the second type as the area type corresponding to any first image area; and, among the first reference number of first image areas, determine the first image area whose corresponding area type is the first type and whose corresponding area type is the second type as the first target area.
[0025] In one possible implementation, the number of the first target areas is at least one, and the determination unit is further configured to divide any first target area into a second reference number of second image areas, determine the number of predicted group objects corresponding to the second reference number of second image areas and the number of standard group objects corresponding to the second reference number of second image areas; determine the area types corresponding to the second reference number of second image areas based on the number of predicted group objects corresponding to the second reference number of second image areas and the number of standard group objects corresponding to the second reference number of second image areas; use a second image area having the same area type as that corresponding to any first target area as a second target area, and the number of the second target areas is at least one; in response to any second target area being composed of a single sample pixel point, use the sample pixel point constituting any second target area as a candidate pixel point in any first target area.
[0026] In one possible implementation, the determination unit is further configured to, in response to any second target area being composed of at least two sample pixels, divide the any second target area into a third reference number of third image areas, determine the number of predicted group objects corresponding to the third reference number of third image areas and the number of standard group objects corresponding to the third reference number of third image areas; determine the region types corresponding to the third reference number of third image areas based on the number of predicted group objects corresponding to the third reference number of third image areas and the number of standard group objects corresponding to the third reference number of third image areas; use a third image area having the same region type as that corresponding to the any second target area as a third target area, and the number of the third target areas is at least one; and in response to any third target area being composed of a single sample pixel, use the sample pixel constituting the any third target area as a candidate pixel in the any first target area.
[0027] In one possible implementation, the determination unit is further used to determine, for any first image area among the first reference number of first image areas, a sub-prediction group density map corresponding to the any first image area in the prediction group density map; determine a sub-standard group density map corresponding to the any first image area in the standard group density map; determine the number of prediction group objects corresponding to the any first image area based on the sub-prediction group density map corresponding to the any first image area; and determine the number of standard group objects corresponding to the any first image area based on the sub-standard group density map corresponding to the any first image area.
[0028] In one possible implementation, the acquisition unit is further used to generate at least one first basic map based on the annotation information of at least one head center point corresponding to the sample image, and the size of any first basic map is the same as the size of the sample image; superimpose the at least one first basic map to obtain a second basic map; and use the target Gaussian kernel to perform Gaussian convolution on the second basic map to obtain a standard group density map corresponding to the sample image.
[0029] In one possible implementation, the target group density map is used to indicate the target group density corresponding to each target pixel point in the target image; the determination unit is used to summarize the target group density corresponding to each target pixel point in the target image to obtain the number of target group objects corresponding to the target image.
[0030] On the other hand, a computer device is provided, comprising a processor and a memory, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement any of the above-mentioned image processing methods.
[0031] On the other hand, a computer-readable storage medium is provided, wherein at least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by a processor to implement any of the above-mentioned image processing methods.
[0032] In another aspect, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any of the above-described image processing methods.
[0033] The technical solutions provided by the embodiments of the present application bring at least the following beneficial effects:
[0034] In an embodiment of the present application, a target image processing model is trained using a first loss function and a second loss function, wherein the first loss function is used to focus on whether the image processing model can accurately predict the predicted population density corresponding to each sample pixel point in the sample image, and the second loss function is used to focus on whether the image processing model can accurately predict the number of predicted population objects corresponding to each first image region in the sample image. The information focused on during the training of the target image processing model is relatively rich, which can alleviate the inconsistency between the training target and the final use target of the image processing model. The target population density map obtained from the target image processing model has a high accuracy in determining the number of target population objects, and the target image processing model is used to process the target image with a better processing effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0036] Figure 1 is a schematic diagram of an implementation environment of an image processing method provided in an embodiment of the present application;
[0037] Figure 2 This is a flowchart of an image processing method provided by an embodiment of the present application;
[0038] Figure 3 This is a schematic diagram of an implementation process of calling a target image processing model to process a target image and obtaining a target group density map corresponding to the target image, provided by an embodiment of the present application;
[0039] Figure 4 Schematic diagram of a process for obtaining a convolution feature Pi provided in an embodiment of the present application;
[0040] Figure 5 This is a visualization of an image itself, a real population density map corresponding to the image, and a target population density map corresponding to the image obtained using a target image processing model, provided in an embodiment of the present application;
[0041] Figure 6 This is a flowchart of a method for training a target image processing model provided by an embodiment of the present application;
[0042] Figure 7 is a schematic diagram of a process of determining a second target area based on any first target area provided by an embodiment of the present application;
[0043] Figure 8 is a visualization diagram of a process of determining candidate pixels in a first target area provided by an embodiment of the present application;
[0044] Figure 9 is a schematic diagram of an image processing device provided in an embodiment of the present application;
[0045] Figure 10 is a schematic diagram of an image processing device provided in an embodiment of the present application;
[0046] Figure 11 This is a schematic diagram of the structure of a terminal provided in an embodiment of the present application;
[0047] Figure 12 This is a structural diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0048] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0049] The image processing method provided in the embodiments of this application relates to the field of artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine that can respond in a manner similar to human intelligence. Artificial intelligence is the study of the design principles and implementation methods of various intelligent machines, giving them the capabilities of perception, reasoning, and decision-making.
[0050] Artificial intelligence technology is a comprehensive discipline covering a wide range of fields, encompassing both hardware and software technologies. Basic AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. Artificial intelligence software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning. The image processing methods provided in the embodiments of this application involve computer vision and machine learning technologies.
[0051] Computer vision (CV) technology is the study of how machines can "see." Specifically, it refers to the use of cameras and computers to replace the human eye in identifying and measuring objects, and further image processing to transform the computer-generated images into images more suitable for human observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems that can extract information from images or multidimensional data. Computer vision technologies generally include image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / action recognition, three-dimensional object reconstruction, 3D (three-dimensional) technology, virtual reality, augmented reality, simultaneous localization and mapping, and other technologies. It also includes common biometric recognition technologies such as facial recognition and fingerprint recognition.
[0052] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning through demonstration.
[0053] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, robots, smart medical care, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0054] In an exemplary embodiment, the image processing method provided in the embodiment of the present application is implemented in a blockchain system. The images involved in the image processing method provided in the embodiment of the present application and the image processing results (such as group density maps, the number of group objects, etc.) are all stored on the blockchain in the blockchain system, and the security and reliability of the images and the image processing results are relatively high.
[0055] This application embodiment provides an image processing method, please refer to Figure 1, which shows a schematic diagram of an implementation environment of the image processing method provided by an embodiment of the present application. The implementation environment includes: a terminal 11 and a server 12.
[0056] The image processing method provided in the embodiment of the present application can be applied to both the terminal 11 and the server 12, and the embodiment of the present application does not limit this. For example, in the case where the image processing method provided in the embodiment of the present application is applied to the terminal 11, the target image to be processed can refer to an image stored locally in the terminal 11, or it can refer to an image obtained from the server 12 or other device (such as an image acquisition device). After obtaining the target group density map and the number of target group objects corresponding to the target image, the terminal 11 can visualize the target group density map and the number of target group objects corresponding to the target image, or it can send the target group density map and the number of target group objects corresponding to the target image to the server 12 for storage.
[0057] Exemplarily, in the case where the image processing method provided in the embodiment of the present application is applied to the server 12, the target image to be processed may refer to an image stored locally on the server 12, or may refer to an image obtained from the terminal 11 or other devices (such as image acquisition devices). After obtaining the target group density map and the number of target group objects corresponding to the target image, the server 12 can send the target group density map and the number of target group objects corresponding to the target image to the terminal 11 for visual display.
[0058] In one possible implementation, terminal 11 may be any electronic product that can interact with a user through one or more methods such as a keyboard, touchpad, touch screen, remote control, voice interaction, or handwriting device, such as a PC (Personal Computer), a mobile phone, a smart phone, a PDA (Personal Digital Assistant), a wearable device, a PPC (Pocket PC), a tablet computer, a smart car computer, a smart TV, a smart speaker, etc. Server 12 may be a single server, a server cluster consisting of multiple servers, or a cloud computing service center. Terminal 11 establishes a communication connection with server 12 via a wired or wireless network.
[0059] Those skilled in the art should understand that the above-mentioned terminal 11 and server 12 are only examples. Other existing or future terminals or servers that are applicable to this application should also be included in the scope of protection of this application and are included here by reference.
[0060] based on Figure 1In the implementation environment shown, the embodiment of the present application provides an image processing method, which is applied to an electronic device, which is a terminal or a server. Figure 2 As shown, the image processing method provided in the embodiment of the present application includes the following steps 201 to 203:
[0061] In step 201, a target image processing model and a target image to be processed are obtained. The target image processing model is trained using a first loss function and a second loss function. The first loss function is determined based on the predicted group density corresponding to each sample pixel point in the sample image, and the second loss function is determined based on the predicted number of group objects corresponding to each first image area in the sample image.
[0062] The target image processing model refers to a model used to process an image and output a group density map corresponding to the image. The group density map is used to indicate the distribution of group objects in the image, and the number of group objects corresponding to the image can be determined based on the group density map. The group is composed of group objects. The embodiment of the present application does not limit the type of group objects. For example, the type of group object is human, or the type of group object is animal (such as dog, sheep, etc.). When the type of group object is human, the group refers to a crowd of people. When the type of group object is animal, the group refers to a group of animals.
[0063] The target image processing model refers to a trained image processing model. Exemplarily, the target image processing model can be pre-trained and stored in the electronic device. In this case, the target image processing model is obtained by directly retrieving it from the storage. Alternatively, the target image processing model can be obtained through real-time training. Regardless of the method, the target image processing model must be trained.
[0064] In an embodiment of the present application, the target image processing model is trained using a first loss function and a second loss function, wherein the first loss function is determined based on the predicted group density corresponding to each sample pixel point in the sample image, and the second loss function is determined based on the predicted number of group objects corresponding to each first image area in the sample image.
[0065] A sample image refers to an image with head center point annotated information used in the process of training a target image processing model. A first loss function, determined based on the predicted population density corresponding to each sample pixel point in the sample image, is used to focus on whether the image processing model can accurately predict the population density corresponding to the sample pixel point. Each first image region in the sample image refers to an image region obtained by dividing the sample image. A second loss function, determined based on the predicted number of group objects corresponding to each first image region in the sample image, is used to focus on whether the image processing model can accurately predict the number of group objects corresponding to the region in the sample image.
[0066] In the process of training the target image processing model using the first loss function and the second loss function, we not only focus on whether the image processing model can accurately predict the group density corresponding to the sample pixel points, but also focus on whether the image processing model can accurately predict the number of group objects corresponding to the area in the sample image. Based on this, the process of training the target image processing model pays attention to richer information, which is conducive to alleviating the inconsistency between the training goal of minimizing the loss function and the final use goal of accurately predicting the group density map that can accurately determine the number of group objects corresponding to the image. The group density map obtained by the trained target image processing model has higher reliability, thereby improving the accuracy of the number of group objects corresponding to the image determined by the group density map. The process of training the target image processing model using the first loss function and the second loss function will be in Figure 6 The embodiment shown is described in detail and will not be described here in detail.
[0067] A target image refers to an image that needs to be processed by an electronic device to infer the number of corresponding group objects. In an exemplary embodiment, the target image may include one or more group objects, and the electronic device may process the target image to determine the number of group objects contained in the target image. In other embodiments, the target image may not include any group objects, and after processing, the electronic device may determine that the number of group objects is zero. The electronic device may obtain the target image to be processed in various ways, and the electronic device may be a terminal or a server.
[0068] In an exemplary embodiment, when the electronic device is a terminal, the ways in which the terminal obtains the target image to be processed include but are not limited to the following: the terminal uses an image acquisition device to acquire the target image to be processed; or, the terminal downloads an image from a target website as the target image to be processed; or, the terminal extracts an image from an image database as the target image to be processed; or, the terminal responds to an image import operation and acquires the imported image as the target image to be processed.
[0069] In an exemplary embodiment, when the electronic device is a server, the ways in which the server obtains the target image to be processed include but are not limited to the following: the server receives the image captured and sent by the terminal using the image acquisition device as the target image to be processed; or, the server downloads the image from the target website as the target image to be processed; or, the server extracts the image from the image database as the target image to be processed.
[0070] It should be noted that the above only provides several possible methods for obtaining the target image to be processed. Of course, the terminal or server can also obtain the target image to be processed through other methods. The embodiment of the present application does not limit the method for obtaining the target image to be processed.
[0071] The embodiment of the present application does not limit the source scene of the target image to be processed.
[0072] In step 202, a target image processing model is called to process the target image to obtain a target group density map corresponding to the target image.
[0073] After obtaining the target image processing model and the target image, the target image processing model is called to process the target image to obtain a target group density map corresponding to the target image. The target group density map is used to indicate the distribution of group objects in the target image. For example, the distribution of group objects in the target image is represented by the target group density corresponding to each target pixel in the target image. In other words, the target group density map is used to indicate the target group density corresponding to each target pixel in the target image. Group density is used to indicate the average number of group objects at the location of a unit pixel. In other words, the target group density corresponding to any target pixel indicates the average number of group objects at the location of that target pixel.
[0074] Exemplarily, the size of the target group density map is the same as the size of the target image, so that the target group density corresponding to each target pixel in the target image can be directly determined based on the target group density map. In exemplary embodiments, the target image and the target group density map in the embodiments of the present application are both two-dimensional. The size of the target group density map being the same as the size of the target image means that the width of the target group density map is the same as the width of the target image, and the height of the target group density map is the same as the height of the target image.
[0075] The embodiment of the present application does not limit the specific implementation method of calling the target image processing model to process the target image and obtain the target group density corresponding to the target image, which may vary depending on the model structure of the target image processing model.
[0076] In one possible implementation, calling a target image processing model to process a target image and obtaining a target population density map corresponding to the target image is accomplished by: calling the target image processing model to extract features from the target image to obtain target convolution features corresponding to the target image; and performing prediction processing on the target convolution features to obtain a target population density map corresponding to the target image. In an exemplary embodiment, the feature extraction process for the target image can be performed by a feature extraction model within the target image processing model, and the prediction processing process for the target convolution features can be performed by a prediction processing model within the target image processing model. This embodiment of the present application does not limit the model structures of the feature extraction model and the prediction processing model.
[0077] Exemplarily, the feature extraction model can be regarded as the backbone network model of the image processing model. The embodiment of the present application does not limit the model structure of the feature extraction model, as long as it can extract features from the image. Exemplarily, the model structure of the feature extraction model is a U-shaped structure; or, the model structure of the feature extraction model is a NASNet (Neural Architecture Search Network) structure; or, the model structure of the feature extraction model is a network structure obtained by directly using a neural network structure search method on the corresponding task. The prediction processing model is used to predict the group density map corresponding to the image based on the target convolution feature. The model structure of the prediction processing model can be a relatively simple model structure. Exemplarily, the prediction processing model is composed of at least one convolutional layer connected in sequence.
[0078] The image processing method provided in the embodiment of the present application can be regarded as a group density estimation method based on deep learning technology. The group density estimation method based on deep learning technology generally takes a single image as input, extracts image features through an image processing model, and then obtains a group density map based on the image features. The group density estimation task usually requires both contextual features with high semantic information and local detail information, so the mainstream network model of the image processing model usually uses a network structure that first downsamples and then upsamples to obtain high-resolution features with both high-level semantic information and detail information, and finally performs a prediction process to output a group density map. In the network structure that first downsamples and then upsamples, a jump link is introduced to introduce detail information for the upsampling process.
[0079] Next, combine Figure 3 , introduces a method of calling a target image processing model to process a target image and obtain a target population density map corresponding to the target image. In the process of calling the target image processing model to process the target image, the downsampling process is first performed, then the upsampling process is performed, and finally the prediction process is performed. Next, combined with Figure 3 , these three processing processes are introduced respectively.
[0080] The downsampling process is as follows: perform a first convolution on the target image to obtain a first convolution feature V1; perform a first pooling on the first convolution feature V1 to obtain a first pooling feature; perform a second convolution on the first pooling feature to obtain a second convolution feature V2; perform a second pooling on the second convolution feature V2 to obtain a second pooling feature; perform a third convolution on the second pooling feature to obtain a third convolution feature V3; perform a third pooling on the third convolution feature V3 to obtain a third pooling feature; perform a fourth convolution on the third pooling feature to obtain a fourth convolution feature V4; perform a fourth pooling on the fourth convolution feature V4 to obtain a fourth pooling feature; perform a fifth convolution on the fourth pooling feature to obtain a fifth convolution feature V5. At this point, the downsampling process is completed.
[0081] During the above-mentioned downsampling process, the sizes of the convolution kernels used in different convolution processes may be the same or different, and the embodiments of the present application do not limit this. For example, the sizes of the convolution kernels used in different convolution processes are all 3×3, but the numbers of convolution kernels used in different convolution processes are different. For example, during the downsampling process, the downsampling steps are {1, 2, 4, 8, 16}, and the convolution features obtained are {V1, V2, V3, V4, V5}. That is, the size of the first convolution feature V1 is the same as the size of the target image, the size of the second convolution feature V2 is 1 / 2 of the size of the target image, the size of the third convolution feature V3 is 1 / 4 of the size of the target image; the size of the fourth convolution feature V4 is 1 / 8 of the size of the target image; the size of the fifth convolution feature V5 is 1 / 16 of the size of the target image.
[0082] It should be noted that the convolution feature is a two-dimensional feature, and the size of the convolution feature is 1 / a of the size of the target image (a is a positive number not less than 1), which means that the width of the convolution feature is 1 / a of the width of the target image, and the height of the convolution feature is 1 / a of the height of the target image.
[0083] In the above-mentioned downsampling process, the pooling methods used in different pooling processes may be the same or different. For example, the pooling methods used in different pooling processes are all maximum pooling.
[0084] Upsampling process: perform the sixth convolution on the fifth convolution feature V5 to obtain the sixth convolution feature P5; upsample the sixth convolution feature P5 to the same size as the fourth convolution feature V4, concatenate the upsampled sixth convolution feature with the fourth convolution feature V4 to obtain the first concatenation feature, perform the seventh convolution on the first concatenation feature to obtain the seventh convolution feature P4; upsample the seventh convolution feature P4 to the same size as the third convolution feature V3, concatenate the upsampled seventh convolution feature with the third convolution feature V3 to obtain the second concatenation feature, and perform the seventh convolution on the second concatenation feature. Perform the eighth convolution process to obtain the eighth convolution feature P3; upsample the eighth convolution feature P3 to the same size as the second convolution feature V2, splice the upsampled eighth convolution feature with the second convolution feature V2 to obtain a third splicing feature, perform the ninth convolution process on the third splicing feature to obtain a ninth convolution feature P2; upsample the ninth convolution feature P2 to the same size as the first convolution feature V1, splice the upsampled ninth convolution feature with the first convolution feature V1 to obtain a fourth splicing feature, perform target convolution process on the fourth splicing feature to obtain a target convolution feature P1.
[0085] In the process of the above-mentioned upsampling processing, the sizes of the convolution kernels used in different convolution processing processes may be the same or different, and the embodiments of the present application do not limit this. Exemplarily, in the convolution processing process, in addition to using the convolution kernel for convolution, operations such as activation using an activation function may also be performed, and the embodiments of the present application do not limit this. Exemplarily, the upsampling methods used in different upsampling processes may be the same or different. Exemplarily, the upsampling methods used in different upsampling processes are all the nearest neighbor interpolation method, or the upsampling methods used in different upsampling processes are all bilinear interpolation methods, etc. Exemplarily, the two features to be spliced have the same size, and the embodiments of the present application do not limit the method of splicing the two features. Exemplarily, the method of splicing the two features is to splice the two features in the channel dimension.
[0086] For example, the downsampling process can be regarded as an encoding process, and the upsampling process can be regarded as a decoding process. In the decoding process, except that P5 is obtained by simply convolving V5, any convolution feature Pi (i=1, 2, 3 or 4) in P1, P2, P3 and P4 is obtained based on the two features Vi and Pi+1. The acquisition process of the convolution feature Pi can be as follows: Figure 4As shown in the figure. First, the convolution feature Pi+1 is upsampled to the same size as Vi using methods such as the nearest neighbor interpolation method. The upsampled feature is then concatenated with Vi to obtain a concatenated feature. The concatenated feature is then subjected to two consecutive 3×3 convolutions (Conv3) and ReLU (Rectified Linear Unit) activations to obtain the convolution feature Pi. The two consecutive 3×3 convolutions and ReLU activations on the concatenated feature are both included in the convolution process of the concatenated feature.
[0087] During the upsampling process, the size of the sixth convolution feature P5 is the same as the size of the fifth convolution feature V5; the size of the seventh convolution feature P4 is the same as the size of the fourth convolution feature V4; the size of the eighth convolution feature P3 is the same as the size of the third convolution feature V3; the size of the ninth convolution feature P2 is the same as the size of the second convolution feature V2; and the size of the target convolution feature P1 is the same as the size of the first convolution feature V1. If the size of the first convolution feature V1 is the same as the size of the target image, the size of the target convolution feature P1 is also the same as the size of the target image.
[0088] exist Figure 3 In the process shown, the target convolution features are predicted by sequentially processing them using two convolutional layers to obtain a target population density map corresponding to the target image. The first convolutional layer has a 3×3 kernel size, denoted as Conv3; the second convolutional layer has a 1×1 kernel size, denoted as Conv1. For example, the target population density map has the same size as the target image.
[0089] It should be noted that the above Figure 3 The illustrated implementation of calling the target image processing model to process the target image and obtain the target population density map corresponding to the target image is merely an example, and the embodiments of the present application are not limited thereto. The process of calling the target image processing model to process the target image and obtain the target population density map corresponding to the target image can be flexibly adjusted based on the model structure of the target image processing model.
[0090] In step 203 , the number of target group objects corresponding to the target image is determined based on the target group density map. The number of target group objects corresponding to the target image is used to indicate the number of group objects contained in the target image.
[0091] The target population density map indicates the target population density corresponding to each target pixel in the target image. The target population density corresponding to any target pixel indicates the average number of group objects at that pixel's location. Based on this, the target population density map can be used to directly determine the number of group objects contained in the target image, i.e., the number of target group objects corresponding to the target image.
[0092] In one possible implementation, based on the target group density map, the process of determining the number of target group objects corresponding to the target image is as follows: summarizing the target group densities corresponding to each target pixel point in the target image to obtain the number of target group objects corresponding to the target image.
[0093] In an exemplary embodiment, the target group density corresponding to each target pixel point in the target image is summarized and processed to obtain the number of target group objects corresponding to the target image. The process is as follows: the target group density corresponding to each target pixel point in the target image is added up to obtain the total group density; based on the total group density, the number of target group objects corresponding to the target image is determined.
[0094] In one possible implementation, the number of target group objects corresponding to a target image is determined based on the total group density by directly using the total group density as the number of target group objects corresponding to the target image. In this case, the value of the number of target group objects may be a decimal. In another possible implementation, the number of target group objects corresponding to a target image is determined based on the total group density by determining an integer value corresponding to the total group density and using this integer value as the number of target group objects corresponding to the target image.
[0095] The embodiments of the present application do not limit the method for determining the integer value corresponding to the total population density. For example, the method for determining the integer value corresponding to the total population density is rounding up; or, the method for determining the integer value corresponding to the total population density is rounding up; or, the method for determining the integer value corresponding to the total population density is rounding down.
[0096] The image processing method provided by the embodiment of the present application can automatically infer the total number of group objects corresponding to an image, playing an important role in fields such as public safety. The image processing method provided by the present application can be applied to any group object number counting scenario.
[0097] For example, embodiments of the present application can be applied on an open image processing platform, taking a single image as input and outputting a group density map corresponding to the image. The group density map can then be used to determine the total number of group objects corresponding to the image, and the group density of each area in the image can also be determined. The number of group objects is calculated by counting the number of head center points of the group objects in the image.
[0098] For example, for any image, the image itself, the real population density map corresponding to the image, and the visualization map of the target population density map corresponding to the image obtained by the target image processing model are respectively as follows: Figure 5 (1) Figure 5 (2) and Figure 5 As shown in (3) in the figure. In the real group density map and the target group density map, different colors are used to represent different group density values. For a certain area in the image, the closer the color in the area is to the background color, the lower the group density in the area; the greater the difference between the color in the area and the background color, the higher the group density in the area. Figure 5 From (2) in the figure, we can see that the number of group objects determined based on the real group density map is 553; Figure 5 From (3) in the figure, we can see that the number of group objects determined based on the target group density map is 551.
[0099] In an embodiment of the present application, a target image processing model is trained using a first loss function and a second loss function, wherein the first loss function is used to focus on whether the image processing model can accurately predict the predicted population density corresponding to each sample pixel point in the sample image, and the second loss function is used to focus on whether the image processing model can accurately predict the number of predicted population objects corresponding to each first image region in the sample image. The information focused on during the training of the target image processing model is relatively rich, which can alleviate the inconsistency between the training target and the final use target of the image processing model. The target population density map obtained from the target image processing model has a high accuracy in determining the number of target population objects, and the target image processing model is used to process the target image with a better processing effect.
[0100] Before calling the target image processing model to process the target image, it is necessary to first train the target image processing model. Figure 6 As shown, the method for training to obtain the target image processing model includes the following steps 601 to 606:
[0101] In step 601, a sample image and a standard population density map corresponding to the sample image are obtained. The standard population density map is used to indicate the standard population density corresponding to each sample pixel point in the sample image.
[0102] Sample images refer to images used to train the initial image processing model. For example, sample images include at least one group of objects. In exemplary embodiments, sample images may also include images that do not include group objects to improve the robustness of the image processing model. This embodiment of the present application does not limit the source of the sample images. Furthermore, this embodiment of the present application does not limit the size of the sample images, which can be set based on experience.
[0103] The sample image may refer to an image stored locally in the electronic device, or may refer to an image obtained by the electronic device from other devices through wired or wireless communication, and this is not limited in the embodiments of the present application. After obtaining the sample image, a standard group density map corresponding to the sample image is further obtained. The standard group density map is used to indicate the standard group density corresponding to each sample pixel point in the sample image. The standard group density corresponding to any sample pixel point is used to indicate the true average number of group objects at the location of the any sample pixel point. Exemplarily, the standard group density corresponding to any sample pixel point is represented by the pixel value corresponding to the any sample pixel point in the standard group density map. The standard group density corresponding to each sample pixel point determined according to the standard group density map is the true group density corresponding to the sample pixel point.
[0104] The image processing model is trained using the standard group density map as the "true value" so that the trained image processing model can process sample images and output the standard group density map, or a group density map that is very close to the standard group density map.
[0105] In one possible implementation, the process of obtaining the standard population density map corresponding to the sample image includes the following steps 6011 to 6013:
[0106] Step 6011: Generate at least one first basic image based on the annotation information of at least one head center point corresponding to the sample image. The size of any first basic image is the same as that of the sample image.
[0107] The annotation information for at least one head center point corresponding to the sample image is used to indicate the location of the head center point of at least one group subject included in the sample image within the sample image. The head center point refers to the geometric center of the head. The annotation information for at least one head center point corresponding to the sample image is used to provide supervision information for the training process of the image processing model.
[0108] In an exemplary embodiment, the annotation information for at least one head center point corresponding to the sample image is bound to the sample image. In this case, the annotation information for at least one head center point corresponding to the sample image can be obtained simultaneously with the acquisition of the sample image. In an exemplary embodiment, the annotation information for at least one head center point corresponding to the sample image is obtained by a professional after acquiring the sample image by annotating the positions of the head centers of the group subjects included in the sample image.
[0109] After obtaining the annotation information of at least one head center point corresponding to the sample image, at least one first base map is generated based on the annotation information of the at least one head center point corresponding to the sample image. The size of each first base map is the same as the size of the sample image. A first base map is generated based on the annotation information of each head center point. A first base map generated based on the annotation information of any head center point is a map that only identifies the position of that head center point.
[0110] In one possible implementation, a first basic map is generated based on the annotation information of any head center point by marking the pixel value at the position indicated by the annotation information of any head center point as 1, and marking the pixel values at other positions as 0, to obtain a first basic map.
[0111] For example, a first basic image obtained based on the annotation information of any head center point can reflect whether each pixel point contains any head center point. If it does, the pixel value of the pixel point is 1. If it does not, the pixel value of the pixel point is 0. Here, "1" represents inclusion and "0" represents non-inclusion as an example. Of course, "0" can also be set to represent inclusion and "1" represents non-inclusion, and this embodiment of the application is not limited to this.
[0112] According to the labeling information of each head center point, a first basic map can be generated, thereby obtaining at least one first basic map. It should be noted that the sizes of the different first basic maps are the same as the size of the sample image.
[0113] Step 6012: Overlay at least one first basic map to obtain a second basic map.
[0114] Each first basic map is used to identify the position of a head center point. After obtaining at least one first basic map, the at least one first basic map is superimposed to obtain a second basic map. The second basic map can identify the positions of all head center points.
[0115] For example, assuming that the position of any head center point is x i , then a first basic graph generated based on the annotation information of any head center point can be expressed as δ(xxi ), in δ(xx i ) in the first basic graph, only position x i The pixel value at the position is 1, and the pixel values at other positions are all 0. Assuming that the number of heads in the sample image is N (N is an integer not less than 1), the N first basic maps generated based on the annotation information of the N head center points are superimposed, and the obtained second basic map can be expressed as It can be noted that the number of group objects corresponding to the sample image can be obtained by integrating the second basic graph.
[0116] Step 6013: Perform Gaussian convolution processing on the second basic image using the target Gaussian kernel to obtain a standard population density map corresponding to the sample image.
[0117] In the second basic map, the group density is concentrated at the center point of the head. The electronic device can perform Gaussian convolution processing on the second basic map to disperse the group density to the pixels around the center point of the head, and then determine the true group density corresponding to each sample pixel of the sample image. The contribution value of each head to the density of the surrounding pixels is attenuated according to the Gaussian function, so the second basic map is Gaussian convolution processed using the target Gaussian convolution kernel, and the image obtained after processing is called the standard group density map corresponding to the sample image. Exemplarily, the target Gaussian convolution kernel is normalized, and integrating the standard group density map can also obtain the total number of group objects contained in the sample image. Exemplarily, integrating the standard group density map refers to summing the standard group densities corresponding to each sample pixel indicated by the standard group density map.
[0118] The target Gaussian kernel may refer to a fixed Gaussian kernel or a geometrically adaptive Gaussian kernel, which is not limited in this embodiment of the present application. In the case where the target Gaussian kernel is a fixed Gaussian kernel, the size of the fixed Gaussian kernel's effective area range σ is set based on experience or flexibly adjusted according to the application scenario, which is not limited in this embodiment of the present application, for example, σ = 4. In the case where the target Gaussian kernel is a geometrically adaptive Gaussian kernel, the size of the effective area range σ of the geometrically adaptive Gaussian kernel is adaptively determined based on the average distance between the center point of the head and the center points of the adjacent heads.
[0119] For example, the second basic graph is represented as As an example, the target Gaussian convolution kernel is represented as G σ , then the process of using the target Gaussian convolution kernel to perform Gaussian convolution on the second basic image to obtain the standard population density map corresponding to the sample image can be expressed as D gt =G σ ×H(x), where D gt Refers to the standard population density.
[0120] In step 602, the initial image processing model is called to process the sample image to obtain a predicted population density map corresponding to the sample image. The predicted population density map is used to indicate the predicted population density corresponding to each sample pixel point in the sample image.
[0121] The initial image processing model refers to the image processing model to be trained, and the predicted population density map refers to the population density map obtained by processing the sample image using the initial image processing model. The predicted population density map corresponding to the sample image indicates the predicted population density corresponding to each sample pixel in the sample image. The predicted population density corresponding to any sample pixel indicates the predicted average number of objects in the population at the location of that sample pixel, as predicted by the initial image processing model.
[0122] The process of calling the initial image processing model to process the sample image and obtaining the predicted population density map corresponding to the sample image is shown in Figure 2 Step 202 in the illustrated embodiment will not be described in detail here.
[0123] In step 603, a first loss function is determined based on the predicted population density corresponding to each sample pixel point and the standard population density corresponding to each sample pixel point.
[0124] The standard population density map is used to indicate the standard population density corresponding to each sample pixel in the sample image, and the predicted population density map is used to indicate the predicted population density corresponding to each sample pixel in the sample image. After obtaining the standard population density map and the predicted population density map, a first loss function can be determined based on the predicted population density corresponding to each sample pixel and the standard population density corresponding to each sample pixel. The first loss function is used to reflect the difference between the predicted population density corresponding to each sample pixel and the standard population density corresponding to each sample pixel. The embodiment of the present application does not limit the method for determining the first loss function based on the predicted population density corresponding to each sample pixel and the standard population density corresponding to each sample pixel.
[0125] Exemplarily, based on the predicted population density corresponding to each sample pixel point and the standard population density corresponding to each sample pixel point, the first loss function is determined as follows: calculating the L2 norm loss function between the predicted population density corresponding to each sample pixel point and the standard population density corresponding to each sample pixel point, and using the calculated L2 norm loss function as the first loss function.
[0126] Exemplarily, the process of calculating the L2 norm loss function between the predicted population density corresponding to each sample pixel point and the standard population density corresponding to each sample pixel point is implemented based on Formula 1:
[0127]
[0128] in, represents the L2 norm loss function between the predicted population density corresponding to each sample pixel and the standard population density corresponding to each sample pixel, that is, the first loss function; D gt Represents the standard population density map; D est Represents the predicted population density map; G represents the set of all sample pixels in the sample image; p∈G represents that p is a sample pixel in the set G.
[0129] In step 604 , the sample image is divided into a first reference number of first image regions, and the predicted number of group objects corresponding to the first reference number of first image regions and the standard number of group objects corresponding to the first reference number of first image regions are determined.
[0130] The first image regions refer to image regions obtained by dividing the sample image. After dividing the sample image, a first reference number of first image regions is obtained. This embodiment of the present application does not limit the first reference number. The first reference number is set based on experience or flexibly adjusted according to the application scenario. Exemplarily, the first reference number is 4. Exemplarily, the sizes of different first image regions can be the same or different, and this embodiment of the present application does not limit this. Exemplarily, the sizes of different first image regions are the same. In other words, the sample image is evenly divided into the first reference number of first image regions.
[0131] Since the first reference number of first image regions are obtained by dividing the sample image, there is no overlap between the first reference number of first image regions, and the first reference number of first image regions are combined to form the sample image.
[0132] After dividing the sample image into a first reference number of first image regions, the number of predicted group objects corresponding to each of the first reference number of first image regions and the number of standard group objects corresponding to each of the first reference number of first image regions are determined. In one possible implementation, the method for determining the number of predicted group objects corresponding to any first image region among the first reference number of first image regions is as follows: determining a sub-predicted group density map corresponding to the any first image region in the predicted group density map; determining a sub-standard group density map corresponding to the any first image region in the standard group density map; determining the number of predicted group objects corresponding to the any first image region based on the sub-predicted group density map corresponding to the any first image region; and determining the number of standard group objects corresponding to the any first image region based on the sub-standard group density map corresponding to the any first image region.
[0133] Any first image area is obtained by dividing the sample image, and the size of the predicted group density map is the same as the size of the sample image. The sub-predicted group density map whose position in the predicted group density map is the same as the position of any first image area in the sample image is used as the sub-predicted group density map corresponding to any first image area; the sub-standard group density map whose position in the standard group density map is the same as the position of any first image area in the sample image is used as the sub-standard group density map corresponding to any first image area.
[0134] In one possible implementation, a sub-prediction population density map corresponding to any first image region is used to indicate the predicted population density corresponding to each sample pixel point in the first image region. Based on the sub-prediction population density map corresponding to the first image region, the number of predicted population objects corresponding to the first image region is determined by aggregating the predicted population densities corresponding to each sample pixel point in the first image region to obtain the number of predicted population objects corresponding to the first image region. Exemplarily, aggregating the predicted population densities corresponding to each sample pixel point in the first image region is a process of calculating the sum of the predicted population densities corresponding to each sample pixel point in the first image region.
[0135] In one possible implementation, a sub-standard population density map corresponding to any first image region is used to indicate the standard population density corresponding to each sample pixel point in the first image region. Based on the sub-standard population density map corresponding to the first image region, the number of standard population objects corresponding to the first image region is determined by aggregating the standard population densities corresponding to each sample pixel point in the first image region to obtain the number of standard population objects corresponding to the first image region. Exemplarily, aggregating the standard population densities corresponding to each sample pixel point in the first image region is a process of calculating the sum of the standard population densities corresponding to each sample pixel point in the first image region.
[0136] The embodiment of the present application introduces a method for determining the predicted number of group objects corresponding to any first image area and the standard number of group objects corresponding to any first image area from the perspective of any first image area. According to the method for determining the predicted number of group objects corresponding to any first image area and the standard number of group objects corresponding to any first image area, the predicted number of group objects corresponding to the first reference number of first image areas and the standard number of group objects corresponding to the first reference number of first image areas can be determined, and then step 605 is executed.
[0137] In step 605 , a second loss function is determined based on the predicted number of group objects corresponding to the first reference number of first image regions and the standard number of group objects corresponding to the first reference number of first image regions.
[0138] After determining the predicted number of group objects corresponding to the first reference number of first image regions and the standard number of group objects corresponding to the first reference number of first image regions, a second loss function is determined based on the predicted number of group objects corresponding to the first reference number of first image regions and the standard number of group objects corresponding to the first reference number of first image regions. The second loss function focuses on whether the image processing model can accurately predict the number of group objects corresponding to each of the first image regions.
[0139] In one possible implementation, based on the predicted number of group objects corresponding to the first reference number of first image regions and the standard number of group objects corresponding to the first reference number of first image regions, determining the second loss function includes the following steps 6051 to 6053:
[0140] Step 6051: Determine a first target area in the first reference number of first image areas based on the predicted number of group objects corresponding to the first reference number of first image areas and the standard number of group objects corresponding to the first reference number of first image areas.
[0141] The first target region refers to a first image region that needs to be further processed. In one possible implementation, step 6051 is implemented as follows: for any first image region among the first reference number of first image regions, in response to the predicted number of group objects corresponding to any first image region being greater than the standard number of group objects corresponding to any first image region, the first type is used as the region type corresponding to any first image region; in response to the predicted number of group objects corresponding to any first image region being less than the standard number of group objects corresponding to any first image region, the second type is used as the region type corresponding to any first image region; and among the first reference number of first image regions, the first image region having the first type and the second type as the first target region.
[0142] The first image region corresponding to the first type of region can be considered an over-predicted first image region, and the first image region corresponding to the second type of region can be considered an under-predicted first image region. Both the over-predicted first image region and the under-predicted first image region are first image regions with poor accuracy in predicting the number of group objects. Using the over-predicted first image region and the under-predicted first image region as the first target region and then obtaining a second loss function based on the first target region can further improve the image processing model's accuracy in predicting the number of group objects corresponding to the image region.
[0143] In an exemplary embodiment, for any first image region, the number of predicted group objects corresponding to the region may be the same as the number of standard group objects corresponding to the region. In this case, the third type is used as the region type corresponding to the first image region. The first image region corresponding to the third type can be considered an accurately predicted first image region, and the process of obtaining the second loss function does not need to focus on the accurately predicted first image region.
[0144] Step 6052: Determine candidate pixel points in the first target area.
[0145] After determining the first target region, candidate pixels within the first target region are further identified. Candidate pixels are sample pixels within the first target region that significantly contribute to the poor accuracy of the predicted number of group objects within the first target region. In exemplary embodiments, candidate pixels are referred to as difficult pixels. During image processing model training, paying more attention to candidate pixels can further improve the training effectiveness of the image processing model.
[0146] It should be noted that the number of first target areas may be one or more, and this is not limited in the embodiments of the present application. In the case where there are multiple first target areas, determining candidate pixel points in the first target areas means determining candidate pixel points in each first target area respectively. In one possible implementation, the process of determining candidate pixel points in any first target area includes the following steps 1 to 5:
[0147] Step 1: Divide any first target area into a second reference number of second image areas, and determine the predicted number of group objects corresponding to the second reference number of second image areas and the standard number of group objects corresponding to the second reference number of second image areas.
[0148] There is at least one first target area. The embodiment of the present application takes any first target area as an example to introduce the process of determining the candidate pixel points in any first target area. The second reference number of second image areas is obtained by dividing the any first target area. The embodiment of the present application does not limit the second reference number. The second reference number can be the same as the first reference number or different from the first reference number. Exemplarily, the second reference number and the first reference number are both 4, that is, any first target area is divided into 4 second image areas. Exemplarily, the sizes of different second image areas can be the same or different, and the embodiment of the present application does not limit this. In an exemplary embodiment, the sizes of different second image areas are the same. That is, any first target area is evenly divided into the second reference number of second image areas.
[0149] Since the second reference number of second image areas are obtained by dividing any first target area, there is no overlap between the second reference number of second image areas, and the second reference number of second image areas are combined to form any first target area.
[0150] After obtaining the second reference number of second image areas, the number of predicted group objects corresponding to the second reference number of second image areas and the number of standard group objects corresponding to the second reference number of second image areas are determined. Exemplarily, the process of determining the number of predicted group objects corresponding to any second image area in the second reference number of second image areas and the number of standard group objects corresponding to the any second image area is as follows: determining a sub-predicted group density map corresponding to the any second image area in the sub-predicted group density map corresponding to the any first target area; determining a sub-standard group density map corresponding to the any second image area in the sub-standard group density map corresponding to the any target area; determining the number of predicted group objects corresponding to the any second image area based on the sub-predicted group density map corresponding to the any second image area; and determining the number of standard group objects corresponding to the any second image area based on the sub-standard group density map corresponding to the any second image area.
[0151] The sub-prediction group density map corresponding to any first target area is determined in the prediction group density map corresponding to the sample image, and the sub-standard density map corresponding to any first target area is determined in the standard group density map corresponding to the sample image. In an exemplary embodiment, the size of the sub-prediction group density map corresponding to any first target area is the same as the size of the any first target area, and the method of determining the sub-prediction group density map corresponding to any second image area in the sub-prediction group density map corresponding to any first target area is: dividing the sub-prediction group density map corresponding to any first target area into a second reference number of sub-prediction group density maps according to the same division method as the division method of the any first target area into a second reference number of second image areas; and using the sub-prediction group density map whose position in the sub-prediction group density map corresponding to any first target area is the same as the position of the any second image area in the any first target area as the sub-prediction group density map corresponding to the any first target area.
[0152] For example, assuming that the second reference number is 4, and any second image area is the second image area in the upper left corner of any first target area, then the sub-prediction group density map in the upper left corner of the sub-prediction group density map corresponding to any first target area is used as the sub-prediction group density map corresponding to any second image area.
[0153] It should be noted that the method for determining the sub-standard group density map corresponding to any second image area in the sub-standard group density map corresponding to any first target area can refer to the method for determining the sub-prediction group density map corresponding to any second image area in the sub-prediction group density map corresponding to any first target area, and will not be repeated here.
[0154] In an exemplary embodiment, after dividing any first target area into a second reference number of second image areas, the second reference number of second image areas can be numbered according to their positions in the any first target area, so as to quickly determine the position of a certain second image area in the any target area according to the number.
[0155] The embodiments of the present application do not limit the numbering method, as long as the second image areas at different positions have different numbers. Exemplarily, the numbering method is sequential numbering from left to right and from top to bottom. Exemplarily, the numbering method can also be to use the coordinates of the center point of the second image area as the number of the second image area. Exemplarily, for the case where the second reference number of second image areas are areas of the same size obtained by evenly dividing any first target area, the numbering method can also be numbering according to the number of rows and columns corresponding to the position of the second image area. For example, the number of any second image area can be expressed as (m, n), which indicates that any second image area is at the mth row and nth column position in any first target area.
[0156] Step 2: Determine the region types corresponding to the second reference number of second image regions based on the predicted number of group objects corresponding to the second reference number of second image regions and the standard number of group objects corresponding to the second reference number of second image regions.
[0157] In one possible implementation, step 2 is implemented as follows: for any second image area in the second reference number of second image areas, in response to the predicted number of group objects corresponding to the second image area being greater than the standard number of group objects corresponding to the second image area, the first type is used as the area type corresponding to the second image area; in response to the predicted number of group objects corresponding to the second image area being less than the standard number of group objects corresponding to the second image area, the second type is used as the area type corresponding to the second image area; and in response to the predicted number of group objects corresponding to the second image area being equal to the standard number of group objects corresponding to the second image area, the third type is used as the area type corresponding to the second image area. In this manner, the area types corresponding to the second reference number of second image areas can be determined.
[0158] Step 3: A second image region having the same region type as that of any first target region is used as a second target region, and the number of the second target region is at least one.
[0159] The area type corresponding to any first target area may be the first type or the second type. When the area type corresponding to any first target area is the first type, the second image area corresponding to the area type of the first type in the second reference number of second image areas is used as the second target area. When the area type corresponding to any first target area is the second type, the second image area corresponding to the area type of the second type in the second reference number of second image areas is used as the second target area.
[0160] That is, if any of the first target regions is an over-predicted region, the over-predicted second image region is used as the second target region; if any of the first target regions is an under-predicted region, the under-predicted second image region is used as the second target region. The second target region is a region that requires further processing to determine candidate pixels for providing data basis for the second loss function.
[0161] The second image region can be regarded as a subregion of any target region. For an over-predicted region, the over-predicted subregion in the region is regarded as a region requiring further processing; for an under-predicted region, the under-predicted subregion in the region is regarded as a region requiring further processing.
[0162] For example, in an over-predicted area, the over-predicted sub-area will lead to a large group object counting error; in an under-predicted area, the under-predicted sub-area will lead to a large group object counting error. Therefore, if more optimization is performed on the over-predicted sub-area in the over-predicted area and the under-predicted sub-area in the under-predicted area, the overall group object counting error may be further reduced. More specifically, taking the over-predicted area as an example, if the over-predicted area is evenly divided into four smaller sub-areas, there is at least one over-predicted sub-area, and the over-predicted sub-area contributes more to the overall over-prediction problem and should be further optimized. On the contrary, the under-predicted sub-area and the correctly predicted sub-area help to alleviate the over-prediction problem of the parent area to a certain extent, and are thus ignored in the process of determining candidate pixels.
[0163] There is at least one second target region, and the specific number of second target regions depends on the actual situation. After determining at least one second target region, each second target region is further processed. For any second target region, the processing process is as follows: determine whether the second target region is composed of a single sample pixel point. If the second target region is composed of a single sample pixel point, execute step 4; if the second target region is composed of at least two sample pixels, execute step 5.
[0164] It should be noted that, among the second reference number of second image areas, there may be second image areas of a different area type in the corresponding area class from that corresponding to any first target area. Such second image areas can be regarded as areas where the corresponding number of group objects is accurately predicted. In the subsequent process of determining candidate pixel points, such second image areas are not considered.
[0165] For example, assuming that the area type corresponding to any first target area is the first type, that is, the number of predicted group objects corresponding to any first target area is greater than the number of standard group objects corresponding to any target area, assuming that the second reference number is 4, the process of determining the second target area based on any first target area is as follows: Figure 7 As shown. Divide any first target area into 4 second image areas, and determine the number of predicted group objects corresponding to each of the 4 second image areas according to the sub-prediction group density map corresponding to any first target area, and then Figure 7 It can be seen that the predicted numbers of group objects corresponding to the four second image regions located at the upper left corner, upper right corner, lower left corner, and lower right corner are 15.53, 17.51, 12.87, and 15.88, respectively.
[0166] According to the sub-standard group density map corresponding to any first target area, the number of standard group objects corresponding to the four second image areas is determined, and the number of standard group objects corresponding to the four second image areas is determined according to the sub-standard group density map corresponding to the first target area. Figure 7 It can be seen that the number of standard group objects corresponding to the four second image regions located in the upper left corner, upper right corner, lower left corner, and lower right corner is 16.27, 17.91, 12.92, and 11.65, respectively. Based on the predicted number of group objects corresponding to each of the four second image regions and the number of standard group objects corresponding to each of the four second image regions, it can be determined that the region type corresponding to the second image region in the upper left corner is the second type, the region type corresponding to the second image region in the upper right corner is the second type, the region type corresponding to the second image region in the lower left corner is the second type, and the region type corresponding to the second image region in the lower right corner is the first type. Since the region type corresponding to any of the first target regions is the first type, the second image region in the lower right corner is selected as the second target region. After determining the second target region, if the second target region consists of at least two sample pixels, the above process is repeatedly performed based on the second target region.
[0167] Step 4: In response to any second object region being composed of a single sample pixel point, the sample pixel point constituting any second object region is used as a candidate pixel point in any first object region.
[0168] The process of determining candidate pixels in any first target region can be considered a search process for that first target region. When any second target region consists of a single sample pixel, it indicates that the search process for that first target region has reached the pixel level. In this case, the sample pixel constituting that second target region is used as a candidate pixel in that first target region.
[0169] Step 5: In response to any second target area being composed of at least two sample pixels, dividing any second target area into a third reference number of third image areas, determining the number of predicted group objects corresponding to the third reference number of third image areas and the number of standard group objects corresponding to the third reference number of third image areas; based on the number of predicted group objects corresponding to the third reference number of third image areas and the number of standard group objects corresponding to the third reference number of third image areas, determining the area types corresponding to the third reference number of third image areas; taking a third image area having the same area type as that corresponding to any second target area as a third target area, and the number of third target areas is at least one; in response to any third target area being composed of a single sample pixel, taking the sample pixel constituting any third target area as a candidate pixel in any first target area.
[0170] When any second target area is composed of at least two sample pixels, it indicates that further searching is required for the second target area, and the process of continuing the search is step 5. The implementation of step 5 is similar to steps 1 to 4 and will not be repeated here.
[0171] It should be noted that the third reference quantity may be the same as the first reference quantity and / or the second reference quantity, or may be different from both the first reference quantity and the second reference quantity, and this embodiment of the present application is not limited to this.
[0172] For example, during the execution of step 5, if it is found that any third target area is composed of at least two sample pixels, the search for any third target area is continued until the target area is degraded to the pixel level, and the sample pixels constituting the pixel-level target area are used as candidate pixels in any first target area. In step 5, only any third target area is used as an example for explanation. For each third target area in at least one third target area, the processing method is the same as that for any third target area. After the processing of at least one third target area is completed, the candidate pixels determined based on the at least one third target area are the candidate pixels in the any first target area determined based on the any second target area.
[0173] It should be noted that steps 4 and 5 only use any second target region as an example to describe the process of determining candidate pixels in any first target region based on that region. The processing method for each of the at least one second target region is the same as for any second target region. After processing the at least one second target region, the candidate pixels determined based on the at least one second target region are the final candidate pixels determined in the at least one first target region.
[0174] It should be further noted that steps 1 through 5 illustrate the process of determining candidate pixels within any first target region using only one first target region as an example. The same processing is applied to each first target region within the at least one first target region. After processing is completed for the at least one first target region, candidate pixels within the at least one first target region are obtained, and step 6053 is then executed.
[0175] For example, the visualization diagram of the process of determining the candidate pixel points in the first target area is as follows: Figure 8 As shown. Figure 8 Different colors are used to identify different types of regions. Within a first-type region, only the first-type subregions are considered; within a second-type region, only the second-type subregions are considered. For each first-type region and each second-type region, the search process is iteratively performed until the pixel level is reached, resulting in candidate pixels in the first target region.
[0176] For example, taking a certain area as the first type of area (i.e., an over-predicted area), and dividing the area into four sub-areas, the algorithm for searching the area and determining the candidate pixels in the area is as follows:
[0177] Input:ε←the estimated heatmap for an over-estimated region R / / Input: The predicted population density map corresponding to the over-estimated region Rε
[0178] D←the ground-truth heatmap for the over-estimated region R / standard population density map D corresponding to the over-estimated region R
[0179] Output:H←the set of the most hard pixels in R / / Output candidate pixel H in the predicted area R
[0180] 1 / / Initialize H to an empty set
[0181] 2PRA_SEARCH(R,ε,D,H) / / Use (R,ε,D,H) to perform the search process
[0182] / *already at pixel level* / / / Already at pixel level
[0183] 3if |ε|=|D|=1then / / ε and D are only used to indicate the population density corresponding to a pixel point
[0184] 4 H←H∪{ε} / / The union of the pixels indicated by H and ε is used as the new H
[0185] 5 return / / return
[0186] / *evenly divide R into four sub-regions R i * / / / Divide R into 4 sub-regions R i
[0187] 6{R1,R2,R3,R4}←divide(R) / / Divide R into {R1,R2,R3,R4}
[0188] 7{ε1,ε2,ε3,ε4}←divide(ε) / / Divide ε into {ε1,ε2,ε3,ε4}
[0189] 8{D1,D2,D3,D4}←divide(D) / / Divide D into {D1,D2,D3,D4}
[0190] / *determine the over-estimated sub-regions* / / / Determine the over-estimated sub-regions
[0191] 9 for i←1 to 4 do / / the value of i is from 1 to 4
[0192] 10 if sum(ε i )>sum(D i )then / / If based on ε i The number of predicted group objects is greater than that based on D i If the number of standard group objects is determined, proceed to the next step
[0193] 11 PRA_SEARCH(R i ,ε i ,D i ,H i ) / / process with the over-estimated sub-regionR i / / Use (R i ,ε i ,D i ,H i ) performs the search process (i.e., based on the over-predicted R i Perform the search process)
[0194] 12 return H / / return H
[0195] The concept of the above search algorithm is as follows: after the sample image is processed by the initial image processing model to obtain a predicted population density map, the sample image is first divided into four image regions of equal size, and then the difficult samples (i.e., candidate pixels) are iteratively searched for in each region. Specifically, for any of the image regions, if the number of predicted population objects is greater than the number of standard population objects, the search process is repeated for the over-predicted sub-regions in the four sub-regions of the image region until the sub-regions degenerate into pixel-level regions. Conversely, if the number of predicted population objects in a certain image region is less than the number of standard population objects, the search process is repeated for the under-predicted sub-regions in the four sub-regions of the image region until the sub-regions degenerate into pixel-level regions. When the search of all image regions degenerates into pixel-level regions, all the determined candidate pixels are further optimized.
[0196] Step 6053: Determine a second loss function based on the predicted population density corresponding to the candidate pixel point and the standard population density corresponding to the candidate pixel point.
[0197] After determining the candidate pixel points, a second loss function is determined based on the predicted population density corresponding to the candidate pixel points and the standard population density corresponding to the candidate pixel points. For example, if region A is a subregion of region B, then region B is called the parent region of region A. Since the candidate pixel points are determined from each sample pixel point of the sample image while considering whether the parent region is an over-predicted region or an under-predicted region, the predicted population density corresponding to the candidate pixel point can be directly determined based on the predicted population density map corresponding to the sample image, and the standard population density corresponding to the candidate pixel point can be directly determined based on the standard population density map corresponding to the sample image.
[0198] It should be noted that the number of candidate pixels determined in step 6052 may be one or more. If there are more than one candidate pixel, the second loss function is determined based on the predicted population density corresponding to each candidate pixel and the standard population density corresponding to each candidate pixel.
[0199] The embodiments of the present application do not limit the method for determining the second loss function based on the predicted population density corresponding to the candidate pixel and the standard population density corresponding to the candidate pixel. For example, the method for determining the second loss function based on the predicted population density corresponding to the candidate pixel and the standard population density corresponding to the candidate pixel is to calculate the L2 norm loss function between the predicted population density corresponding to the candidate pixel and the standard population density corresponding to the candidate pixel, and use the calculated L2 norm loss function as the second loss function.
[0200] Exemplarily, the process of calculating the L2 norm loss function between the predicted population density corresponding to the candidate pixel and the standard population density corresponding to the candidate pixel is implemented based on Formula 2:
[0201]
[0202] in, represents the L2 norm loss function between the predicted population density corresponding to the candidate pixel and the standard population density corresponding to the candidate pixel, which is also the second loss function; D gt Represents the standard population density map; D est Represents the predicted population density map; H represents the set of all candidate pixel points; p∈H means that p is a candidate pixel point in the set H.
[0203] In step 606, the initial image processing model is trained using the first loss function and the second loss function to obtain a target image processing model.
[0204] After determining the first loss function based on step 603 and the second loss function based on step 605, the initial image processing model is trained using the first loss function and the second loss function. Training the initial image processing model using the first loss function and the second loss function can pay more attention to candidate pixels on the basis of paying attention to all sample pixels, thereby being able to optimize the prediction results of candidate pixels in a targeted manner. Since candidate pixels are sample pixels that cause inaccurate predicted number of group objects corresponding to the region, optimizing the prediction results of candidate pixels in a targeted manner can further improve the accuracy of the predicted number of group objects, thereby alleviating the inconsistency between the training objectives of the image processing model and the ultimate use objectives of the image processing model.
[0205] In one possible implementation, the process of training the initial image processing model using the first loss function and the second loss function to obtain the target image processing model is as follows: determining the target loss function based on the first loss function and the second loss function; and training the initial image processing model using the target loss function to obtain the target image processing model. The embodiment of the present application does not limit the method of determining the target loss function based on the first loss function and the second loss function. For example, the process of determining the target loss function based on the first loss function and the second loss function is implemented based on Formula 3:
[0206]
[0207] in, represents the target loss function; represents the first loss function; Represents the second loss function; γ represents the weight coefficient added to the second loss function. The value of γ is set according to experience or flexibly adjusted according to the application scenario. The embodiment of the present application does not limit this. For example, the value of γ is 1.
[0208] In the process of determining the target loss function based on Formula 3, after all candidate pixels are obtained, all sample pixels and candidate pixels are optimized together. Candidate pixels are given more attention through the weight coefficient γ, so that these candidate pixels can be optimized in a targeted manner.
[0209] It should be noted that the number of sample images used to determine the target loss function can be one or more, and this embodiment of the present application does not limit this.
[0210] After determining the target loss function, the target loss function is used to train the initial image processing model to obtain the target image processing model. Exemplarily, the process of training the initial image processing model using the target loss function to obtain the target image processing model is as follows: using the target loss function to update the parameters of the initial image processing model to obtain a first image processing model; in response to the training process satisfying the termination condition, the first image processing model is used as the target image processing model; in response to the training process not satisfying the termination condition, the sample image and the first image processing model are used to continue to obtain the updated target loss function, and then the updated target loss function is used to update the parameters of the first image processing model to obtain a second image processing model, and so on, until the training process satisfies the termination condition, and the image processing model obtained when the training process satisfies the termination condition is used as the target image processing model.
[0211] The process of continuing to obtain the updated target loss function using the sample image and the first image processing model is described in steps 601 to 606 and will not be repeated here. It should be noted that the sample image based on which the target loss function for updating the first image processing model is obtained may be partially or completely the same as the sample image based on which the target loss function for updating the initial image processing model is obtained, or may be completely different from the sample image based on which the target loss function for updating the initial image processing model is obtained, and this embodiment of the present application is not limited to this.
[0212] In one possible implementation, the training process terminates when one of the following conditions is met: the target loss function converges; the target loss function decreases to a loss function threshold; or the number of updates to the image processing model parameters reaches a threshold. The loss function threshold and the threshold are set empirically or flexibly adjusted based on the application scenario, and are not limited in this embodiment.
[0213] After obtaining the target image processing model, the target image processing model can be used to process the image to be processed to more accurately infer the group density map corresponding to the image to be processed, and then the inferred group density map can be used to more accurately determine the number of group objects corresponding to the image to be processed.
[0214] In the embodiment of the present application, a new loss function is used to train the image processing model to alleviate the problem of inconsistency between the training objective and the final use objective, that is, minimizing the training loss does not necessarily guarantee the highest counting accuracy during reasoning. Specifically, on the basis of the traditional L2 loss function (i.e., the first loss function) determined based on each sample pixel point, an additional pyramid region perception loss function (i.e., the second loss function) from the region level to the pixel level is considered. The region perception loss function is determined based on difficult samples (i.e., candidate pixels) and can give more weight to difficult samples. Difficult samples are iteratively screened out according to whether the number of predicted group objects in the parent area is greater than or less than the number of standard group objects. By considering the region perception loss function, these difficult samples can be further optimized. The introduction of this region perception loss function can well alleviate the problem of inconsistency between the training objective and the final use objective.
[0215] When selecting difficult samples, the over-prediction or under-prediction of the parent region is considered from top to bottom. For over-prediction regions, the over-predicted sub-regions within that region are selected as difficult sub-regions, while for under-prediction regions, the under-predicted sub-regions within that region are selected as difficult sub-regions. This selection process is performed iteratively from top to bottom, starting with the image regions into which the sample image is divided, until all difficult samples are finally filtered out. Considering the prediction results of the parent region each time when filtering difficult samples can alleviate the inconsistency between the closest pixel-level prediction during optimization and the closest full-image group object count during inference, thereby improving model prediction accuracy.
[0216] See also Figure 9 , an embodiment of the present application provides an image processing device, the device comprising:
[0217] An acquisition unit 901 is configured to acquire a target image processing model and a target image to be processed, wherein the target image processing model is trained using a first loss function and a second loss function, wherein the first loss function is determined based on a predicted population density corresponding to each sample pixel point in the sample image, and the second loss function is determined based on a predicted number of population objects corresponding to each first image region in the sample image;
[0218] The calling unit 902 is used to call the target image processing model to process the target image to obtain a target group density map corresponding to the target image;
[0219] The determining unit 903 is configured to determine the number of target group objects corresponding to the target image based on the target group density map. The number of target group objects corresponding to the target image is used to indicate the number of group objects included in the target image.
[0220] In a possible implementation, the acquiring unit 901 is further configured to acquire a sample image and a standard population density map corresponding to the sample image, where the standard population density map indicates the standard population density corresponding to each sample pixel point in the sample image.
[0221] The calling unit 902 is further configured to call the initial image processing model to process the sample image to obtain a predicted population density map corresponding to the sample image, where the predicted population density map indicates the predicted population density corresponding to each sample pixel in the sample image.
[0222] The determining unit 903 is further configured to determine a first loss function based on the predicted population density corresponding to each sample pixel point and the standard population density corresponding to each sample pixel point;
[0223] See also Figure 10 , the device further comprises:
[0224] a dividing unit 904, configured to divide the sample image into a first reference number of first image regions;
[0225] The determining unit 903 is further configured to determine the number of predicted group objects corresponding to the first reference number of first image regions and the number of standard group objects corresponding to the first reference number of first image regions;
[0226] The determining unit 903 is further configured to determine a second loss function based on the predicted number of group objects corresponding to the first reference number of first image regions and the standard number of group objects corresponding to the first reference number of first image regions;
[0227] See also Figure 10 , the device further comprises:
[0228] The training unit 905 is used to train the initial image processing model using the first loss function and the second loss function to obtain a target image processing model.
[0229] In one possible implementation, the determination unit 903 is further used to determine the first target area in the first image area of the first reference number based on the predicted number of group objects corresponding to the first image area of the first reference number and the standard number of group objects corresponding to the first image area of the first reference number; determine the candidate pixel points in the first target area; and determine the second loss function based on the predicted group density corresponding to the candidate pixel points and the standard group density corresponding to the candidate pixel points.
[0230] In one possible implementation, the determination unit 903 is further configured to, for any first image area among the first reference number of first image areas, in response to the predicted number of group objects corresponding to any first image area being greater than the standard number of group objects corresponding to any first image area, use the first type as the area type corresponding to any first image area; in response to the predicted number of group objects corresponding to any first image area being less than the standard number of group objects corresponding to any first image area, use the second type as the area type corresponding to any first image area; and, among the first image areas of the first reference number, determine the first image area whose corresponding area type is the first type and whose corresponding area type is the second type as the first target area.
[0231] In one possible implementation, the number of first target areas is at least one, and the determination unit 903 is further configured to divide any first target area into a second reference number of second image areas, determine the number of predicted group objects corresponding to the second reference number of second image areas and the number of standard group objects corresponding to the second reference number of second image areas; determine the area types corresponding to the second reference number of second image areas based on the number of predicted group objects corresponding to the second reference number of second image areas and the number of standard group objects corresponding to the second reference number of second image areas; use a second image area with the same area type as that corresponding to any first target area as a second target area, and the number of second target areas is at least one; and in response to any second target area being composed of a single sample pixel point, use the sample pixel point constituting any second target area as a candidate pixel point in any first target area.
[0232] In one possible implementation, the determination unit 903 is further configured to, in response to any second target area being composed of at least two sample pixels, divide any second target area into a third reference number of third image areas, determine the number of predicted group objects corresponding to the third reference number of third image areas and the number of standard group objects corresponding to the third reference number of third image areas; determine the region types corresponding to the third reference number of third image areas based on the number of predicted group objects corresponding to the third reference number of third image areas and the number of standard group objects corresponding to the third reference number of third image areas; use a third image area having the same region type as that corresponding to any second target area as a third target area, and the number of third target areas is at least one; and in response to any third target area being composed of a single sample pixel, use the sample pixel constituting any third target area as a candidate pixel in any first target area.
[0233] In one possible implementation, the determination unit 903 is further used to determine, for any first image area in the first reference number of first image areas, a sub-prediction group density map corresponding to any first image area in the prediction group density map; determine a sub-standard group density map corresponding to any first image area in the standard group density map; determine the number of prediction group objects corresponding to any first image area based on the sub-prediction group density map corresponding to any first image area; and determine the number of standard group objects corresponding to any first image area based on the sub-standard group density map corresponding to any first image area.
[0234] In one possible implementation, the acquisition unit 901 is further used to generate at least one first basic map based on the annotation information of at least one head center point corresponding to the sample image, and the size of any first basic map is the same as the size of the sample image; superimpose the at least one first basic map to obtain a second basic map; and use the target Gaussian kernel to perform Gaussian convolution on the second basic map to obtain a standard group density map corresponding to the sample image.
[0235] In one possible implementation, the target group density map is used to indicate the target group density corresponding to each target pixel point in the target image; the determination unit 903 is used to summarize the target group density corresponding to each target pixel point in the target image to obtain the number of target group objects corresponding to the target image.
[0236] In an embodiment of the present application, a target image processing model is trained using a first loss function and a second loss function, wherein the first loss function is used to focus on whether the image processing model can accurately predict the predicted population density corresponding to each sample pixel point in the sample image, and the second loss function is used to focus on whether the image processing model can accurately predict the number of predicted population objects corresponding to each first image region in the sample image. The information focused on during the training of the target image processing model is relatively rich, which can alleviate the inconsistency between the training target and the final use target of the image processing model. The target population density map obtained from the target image processing model has a high accuracy in determining the number of target population objects, and the target image processing model is used to process the target image with a better processing effect.
[0237] It should be noted that the apparatus provided in the above embodiments is merely illustrated by the division of the above functional modules when implementing its functions. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0238] Figure 11 This is a schematic diagram of the structure of a terminal provided in an embodiment of the present application. The terminal may be a smartphone, tablet computer, laptop computer, or desktop computer. The terminal may also be referred to as user equipment, portable terminal, laptop terminal, desktop terminal, or other names.
[0239] Typically, the terminal includes: a processor 1101 and a memory 1102 .
[0240] The processor 1101 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1101 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1101 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1101 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1101 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0241] The memory 1102 may include one or more computer-readable storage media, which may be non-transitory. The memory 1102 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1102 is used to store at least one instruction, which is executed by the processor 1101 to implement the image processing method provided in the method embodiment of the present application.
[0242] In some embodiments, the terminal may optionally include a peripheral device interface 1103 and at least one peripheral device. The processor 1101, memory 1102, and peripheral device interface 1103 may be connected via a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1103 via a bus, signal lines, or circuit boards. Specifically, the peripheral device may include at least one of a radio frequency circuit 1104, a display screen 1105, a camera assembly 1106, an audio circuit 1107, and a power supply 1109.
[0243] The peripheral device interface 1103 can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 1101 and the memory 1102. In some embodiments, the processor 1101, the memory 1102, and the peripheral device interface 1103 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1101, the memory 1102, and the peripheral device interface 1103 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0244] The RF circuit 1104 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1104 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1104 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the RF circuit 1104 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and the like. The RF circuit 1104 can communicate with other terminals via at least one wireless communication protocol. Such wireless communication protocols include, but are not limited to, metropolitan area networks, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1104 may also include circuits related to NFC (Near Field Communication), which is not limited in this application.
[0245] Display screen 1105 is used to display a user interface (UI). This UI may include graphics, text, icons, videos, or any combination thereof. When display screen 1105 is a touchscreen display, it is also capable of collecting touch signals on or above the surface of display screen 1105. These touch signals can be input as control signals to processor 1101 for processing. Display screen 1105 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there can be a single display screen 1105, located on the front panel of the terminal. In other embodiments, there can be at least two display screens 1105, located on different surfaces of the terminal or in a foldable design. In still other embodiments, display screen 1105 can be a flexible display, located on a curved or foldable surface of the terminal. Display screen 1105 can also be configured as a non-rectangular, irregular shape, also known as a special-shaped screen. Display screen 1105 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0246] The camera assembly 1106 is used to capture images or videos. Optionally, the camera assembly 1106 includes a front camera and a rear camera. Typically, the front camera is arranged on the front panel of the terminal, and the rear camera is arranged on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 1106 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.
[0247] The audio circuit 1107 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals and input them into the processor 1101 for processing, or input them into the RF circuit 1104 to achieve voice communication. For the purpose of stereo sound collection or noise reduction, there can be multiple microphones, which are respectively set at different parts of the terminal. The microphone can also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signals from the processor 1101 or the RF circuit 1104 into sound waves. The speaker can be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signals into sound waves audible to humans, but also convert the electrical signals into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 1107 may also include a headphone jack.
[0248] Power supply 1109 is used to power various components in the terminal. Power supply 1109 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 1109 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.
[0249] In some embodiments, the terminal further includes one or more sensors 1110 , including but not limited to: an acceleration sensor 1111 , a gyroscope sensor 1112 , a pressure sensor 1113 , an optical sensor 1115 , and a proximity sensor 1116 .
[0250] The accelerometer 1111 can detect the magnitude of acceleration along the three coordinate axes of the coordinate system established by the terminal. For example, the accelerometer 1111 can be used to detect the components of gravity acceleration along the three coordinate axes. The processor 1101 can control the display screen 1105 to display the user interface in a landscape or portrait view based on the gravity acceleration signal collected by the accelerometer 1111. The accelerometer 1111 can also be used to collect game or user motion data.
[0251] The gyroscope sensor 1112 can detect the terminal's body orientation and rotation angle. It can also work with the accelerometer 1111 to collect the user's 3D movements of the terminal. Based on the data collected by the gyroscope sensor 1112, the processor 1101 can implement the following functions: motion sensing (such as changing the UI based on the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.
[0252] The pressure sensor 1113 can be set in the side frame of the terminal and / or the lower layer of the display screen 1105. When the pressure sensor 1113 is set in the side frame of the terminal, it can detect the user's grip signal of the terminal, and the processor 1101 performs left and right hand recognition or shortcut operations based on the grip signal collected by the pressure sensor 1113. When the pressure sensor 1113 is set in the lower layer of the display screen 1105, the processor 1101 controls the operable controls on the UI interface based on the user's pressure operation on the display screen 1105. Operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0253] Optical sensor 1115 is used to collect ambient light intensity. In one embodiment, processor 1101 can control the display brightness of display screen 1105 based on the ambient light intensity collected by optical sensor 1115. Specifically, when the ambient light intensity is high, the display brightness of display screen 1105 is increased; when the ambient light intensity is low, the display brightness of display screen 1105 is decreased. In another embodiment, processor 1101 can also dynamically adjust the shooting parameters of camera assembly 1106 based on the ambient light intensity collected by optical sensor 1115.
[0254] Proximity sensor 1116, also known as a distance sensor, is typically located on the front panel of the terminal. Proximity sensor 1116 is used to detect the distance between the user and the front of the terminal. In one embodiment, when proximity sensor 1116 detects that the distance between the user and the front of the terminal is gradually decreasing, processor 1101 controls display screen 1105 to switch from the screen-on state to the screen-off state. When proximity sensor 1116 detects that the distance between the user and the front of the terminal is gradually increasing, processor 1101 controls display screen 1105 to switch from the screen-off state to the screen-on state.
[0255] Those skilled in the art will understand that Figure 11 The structure shown in the figure does not constitute a limitation on the terminal, and may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.
[0256] Figure 12This is a schematic diagram of the structure of a server provided in an embodiment of the present application. The server may have relatively large differences due to different configurations or performances, and may include one or more processors (Central Processing Units, CPU) 1201 and one or more memories 1202, wherein the one or more memories 1202 store at least one computer program, and the at least one computer program is loaded and executed by the one or more processors 1201 to implement the image processing methods provided in the above-mentioned various method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input and output interface for input and output. The server may also include other components for implementing device functions, which will not be described in detail here.
[0257] In an exemplary embodiment, a computer device is further provided, comprising a processor and a memory, wherein the memory stores at least one computer program, which is loaded and executed by one or more processors to implement any of the above-mentioned image processing methods.
[0258] In an exemplary embodiment, a computer-readable storage medium is further provided, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor of a computer device to implement any of the above-mentioned image processing methods.
[0259] In one possible implementation, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, and the like.
[0260] In an exemplary embodiment, a computer program product or computer program is also provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any of the above-described image processing methods.
[0261] It should be noted that the terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments of the application described herein can be implemented in an order other than those illustrated or described herein. The implementations described in the above exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with certain aspects of the present application as detailed in the appended claims.
[0262] It should be understood that the term "plurality" used herein refers to two or more. "And / or" describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates an "or" relationship between the associated objects.
[0263] The above description is merely an exemplary embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. An image processing method, characterized in that: The method comprises: Obtaining a target image processing model and a target image to be processed, the target image processing model being trained using a first loss function and a second loss function, the first loss function being determined based on predicted population densities corresponding to respective sample pixels in a sample image, and the second loss function being determined based on predicted population densities corresponding to candidate pixels in a first target region and standard population densities corresponding to the candidate pixels, the first target region being a first image region of a first reference quantity in which a corresponding predicted number of population objects is inaccurate, the candidate pixels being sample pixels that cause the predicted number of population objects corresponding to the first target region to be inaccurate, and the first image region of the first reference quantity being obtained by dividing the sample image; Calling the target image processing model to process the target image to obtain a target group density map corresponding to the target image; The number of target group objects corresponding to the target image is determined based on the target group density map, where the number of target group objects corresponding to the target image is used to indicate the number of group objects contained in the target image.
2. The method according to claim 1, characterized in that Before obtaining the target image processing model and the target image to be processed, the method further includes: Acquire the sample image and a standard population density map corresponding to the sample image, wherein the standard population density map is used to indicate the standard population density corresponding to each sample pixel point in the sample image; Invoking an initial image processing model to process the sample image to obtain a predicted population density map corresponding to the sample image, wherein the predicted population density map is used to indicate the predicted population density corresponding to each sample pixel point in the sample image; Determining the first loss function based on the predicted population density corresponding to each sample pixel point and the standard population density corresponding to each sample pixel point; Dividing the sample image into a first reference number of first image regions, and determining a predicted number of group objects corresponding to each of the first reference number of first image regions and a standard number of group objects corresponding to each of the first reference number of first image regions; determining the first target area in the first reference number of first image areas based on the predicted number of group objects corresponding to the first reference number of first image areas and the standard number of group objects corresponding to the first reference number of first image areas; Determining the candidate pixel points in the first target area; determining the second loss function based on the predicted population density corresponding to the candidate pixel and the standard population density corresponding to the candidate pixel; The initial image processing model is trained using the first loss function and the second loss function to obtain the target image processing model.
3. The method according to claim 2, characterized in that The determining, based on the predicted number of group objects respectively corresponding to the first reference number of first image areas and the standard number of group objects respectively corresponding to the first reference number of first image areas, the first target area in the first reference number of first image areas comprises: For any first image region among the first reference number of first image regions, in response to a predicted number of group objects corresponding to the any first image region being greater than a standard number of group objects corresponding to the any first image region, using the first type as the region type corresponding to the any first image region; In response to the predicted number of group objects corresponding to any of the first image regions being less than the standard number of group objects corresponding to any of the first image regions, using the second type as the region type corresponding to any of the first image regions; Among the first reference number of first image regions, first image regions corresponding to the first type and corresponding to the second type are determined as the first target regions.
4. The method according to claim 3, characterized in that The number of the first target area is at least one, and determining the candidate pixel point in the first target area includes: Divide any first target area into a second reference number of second image areas, and determine the predicted number of group objects corresponding to the second reference number of second image areas and the standard number of group objects corresponding to the second reference number of second image areas; determining, based on the predicted number of group objects respectively corresponding to the second reference number of second image areas and the standard number of group objects respectively corresponding to the second reference number of second image areas, the area types respectively corresponding to the second reference number of second image areas; taking a second image region having the same region type as that corresponding to any of the first target regions as a second target region, where the number of the second target region is at least one; In response to any second object region being composed of a single sample pixel point, the sample pixel point constituting any second object region is used as a candidate pixel point in any first object region.
5. The method according to claim 4, characterized in that The method further comprises: In response to any second target area being composed of at least two sample pixels, dividing the any second target area into a third reference number of third image areas, and determining a predicted number of group objects corresponding to each of the third reference number of third image areas and a standard number of group objects corresponding to each of the third reference number of third image areas; determining, based on the predicted number of group objects respectively corresponding to the third reference number of third image areas and the standard number of group objects respectively corresponding to the third reference number of third image areas, the area types respectively corresponding to the third reference number of third image areas; taking a third image region having the same region type as that corresponding to any second target region as a third target region, where the number of the third target region is at least one; In response to any third object region being composed of a single sample pixel point, the sample pixel point constituting any third object region is used as a candidate pixel point in any first object region.
6. The method according to any one of claims 2 to 5, characterized in that: The determining the predicted number of group objects corresponding to the first reference number of first image regions and the standard number of group objects corresponding to the first reference number of first image regions includes: For any first image region among the first reference number of first image regions, determining a sub-prediction population density map corresponding to the first image region in the prediction population density map; and determining a sub-standard population density map corresponding to the first image region in the standard population density map; determining the number of predicted group objects corresponding to any first image region based on the sub-prediction group density map corresponding to any first image region; The number of standard group objects corresponding to any first image region is determined based on the sub-standard group density map corresponding to any first image region.
7. The method according to any one of claims 2 to 5, characterized in that: The method for obtaining the standard population density map corresponding to the sample image includes: generating at least one first basic image based on the annotation information of at least one head center point corresponding to the sample image, wherein the size of any first basic image is the same as the size of the sample image; performing a superposition process on the at least one first basic map to obtain a second basic map; The second basic image is subjected to Gaussian convolution processing using a target Gaussian kernel to obtain a standard population density map corresponding to the sample image.
8. The method according to any one of claims 1 to 5, characterized in that: The target group density map is used to indicate the target group density corresponding to each target pixel point in the target image; and determining the number of target group objects corresponding to the target image based on the target group density map includes: The target group densities corresponding to the target pixels in the target image are aggregated to obtain the number of target group objects corresponding to the target image.
9. An image processing device, characterized in that: The device comprises: an acquisition unit, configured to acquire a target image processing model and a target image to be processed, the target image processing model being trained using a first loss function and a second loss function, the first loss function being determined based on a predicted population density corresponding to each sample pixel in a sample image, and the second loss function being determined based on a predicted population density corresponding to a candidate pixel in a first target region and a standard population density corresponding to the candidate pixel, the first target region being a first image region of a first reference quantity in which a corresponding predicted number of group objects is inaccurate, the candidate pixel being a sample pixel that causes an inaccurate predicted number of group objects corresponding to the first target region, and the first image region of the first reference quantity being obtained by dividing the sample image; a calling unit, configured to call the target image processing model to process the target image, and obtain a target group density map corresponding to the target image; A determining unit is configured to determine the number of target group objects corresponding to the target image based on the target group density map, where the number of target group objects corresponding to the target image is used to indicate the number of group objects contained in the target image.
10. The device according to claim 9, characterized in that The acquisition unit is further configured to acquire the sample image and a standard population density map corresponding to the sample image, wherein the standard population density map is configured to indicate the standard population density corresponding to each sample pixel point in the sample image; The calling unit is further configured to call the initial image processing model to process the sample image to obtain a predicted population density map corresponding to the sample image, wherein the predicted population density map is used to indicate the predicted population density corresponding to each sample pixel point in the sample image; The determining unit is further configured to determine the first loss function based on the predicted population density corresponding to each sample pixel point and the standard population density corresponding to each sample pixel point; The device further comprises: a dividing unit, configured to divide the sample image into a first reference number of first image regions; The determining unit is further configured to determine the predicted number of group objects corresponding to the first reference number of first image areas and the standard number of group objects corresponding to the first reference number of first image areas; determine the first target area in the first reference number of first image areas based on the predicted number of group objects corresponding to the first reference number of first image areas and the standard number of group objects corresponding to the first reference number of first image areas; determine the candidate pixels in the first target area; and determine the second loss function based on the predicted group density corresponding to the candidate pixels and the standard group density corresponding to the candidate pixels. The device further comprises: A training unit is used to train the initial image processing model using the first loss function and the second loss function to obtain the target image processing model.
11. The device according to claim 10, characterized in that The determining unit is further configured to, for any first image region among the first reference number of first image regions, in response to a predicted number of group objects corresponding to the any first image region being greater than a standard number of group objects corresponding to the any first image region, set the first type as the region type corresponding to the any first image region; and in response to a predicted number of group objects corresponding to the any first image region being less than the standard number of group objects corresponding to the any first image region, set the second type as the region type corresponding to the any first image region; Among the first reference number of first image regions, first image regions corresponding to the first type and corresponding to the second type are determined as the first target regions.
12. The device according to claim 11, characterized in that The number of the first target areas is at least one, and the determining unit is further configured to divide any first target area into a second reference number of second image areas, determine the predicted number of group objects corresponding to each of the second reference number of second image areas and the standard number of group objects corresponding to each of the second reference number of second image areas; and determine the area types corresponding to each of the second reference number of second image areas based on the predicted number of group objects corresponding to each of the second reference number of second image areas and the standard number of group objects corresponding to each of the second reference number of second image areas. taking a second image region having the same region type as that corresponding to any of the first target regions as a second target region, where the number of the second target region is at least one; In response to any second object region being composed of a single sample pixel point, the sample pixel point constituting any second object region is used as a candidate pixel point in any first object region.
13. The device according to claim 12, characterized in that The determining unit is further configured to, in response to any second target area being composed of at least two sample pixels, divide the any second target area into a third reference number of third image areas, determine the number of predicted group objects corresponding to each of the third reference number of third image areas and the number of standard group objects corresponding to each of the third reference number of third image areas; and determine the area types corresponding to each of the third reference number of third image areas based on the number of predicted group objects corresponding to each of the third reference number of third image areas and the number of standard group objects corresponding to each of the third reference number of third image areas. taking a third image region having the same region type as that corresponding to any second target region as a third target region, where the number of the third target region is at least one; In response to any third object region being composed of a single sample pixel point, the sample pixel point constituting any third object region is used as a candidate pixel point in any first object region.
14. The device according to any one of claims 10 to 13, characterized in that: The determination unit is further configured to determine, for any first image area among the first reference number of first image areas, a sub-prediction group density map corresponding to the any first image area in the prediction group density map; determine a sub-standard group density map corresponding to the any first image area in the standard group density map; determine the number of prediction group objects corresponding to the any first image area based on the sub-prediction group density map corresponding to the any first image area; and determine the number of standard group objects corresponding to the any first image area based on the sub-standard group density map corresponding to the any first image area.
15. The device according to any one of claims 10 to 13, characterized in that: The acquisition unit is further configured to generate at least one first basic image based on the annotation information of at least one head center point corresponding to the sample image, wherein the size of any first basic image is the same as the size of the sample image; The at least one first basic map is overlaid to obtain a second basic map; and the second basic map is Gaussian convolution processed using a target Gaussian kernel to obtain a standard population density map corresponding to the sample image.
16. The device according to any one of claims 9 to 13, characterized in that: The target group density map is used to indicate the target group density corresponding to each target pixel point in the target image; the determination unit is used to summarize the target group density corresponding to each target pixel point in the target image to obtain the number of target group objects corresponding to the target image.
17. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the image processing method according to any one of claims 1 to 8.
18. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one computer program, and the at least one computer program is loaded and executed by a processor to implement the image processing method according to any one of claims 1 to 8.
19. A computer program product, characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the image processing method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Method and apparatus for generating predictive models
CN107624189A
Dense crowd counting algorithm based on cascaded high-resolution convolutional neural network
CN111460912A