Method and system for counting number of tourists based on multi-scale feature fusion deep snake

By introducing multi-scale feature fusion technology into the Deep Snake algorithm, the receptive field and enhance shallow features are solved, and the accuracy and efficiency of traditional tourist number statistics methods in scenic spots without ticket purchase/ticket check are achieved, and high-precision automatic tourist number statistics are achieved.

CN119992454AActive Publication Date: 2025-05-13HUNAN SANY IND VOCATIONAL & TECH COLLEGE
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510071972.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

The traditional tourist number statistics method is time-consuming and labor-intensive and has a large error, especially in scenic spots where ticket purchases/checks are not required.

Method used

The algorithm based on multi-scale features fusion Deep Snake is used to expand the receptive field and enhance shallow features to achieve fine segmentation and detection of portrait features in the image, thereby counting the number of tourists.

Benefits of technology

It improves the accuracy and efficiency of tourist numbers, and can realize automatic statistics in scenic spots without purchasing or checking tickets, overcoming the shortcomings of traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992454A_ABST
    Figure CN119992454A_ABST
Patent Text Reader

Abstract

The invention discloses a tourist number statistical method and system based on multi-scale feature fusion Deep Snake, and the method comprises the steps: photographing a tourist image according to a camera, employing a Center Net target detection network to detect a portrait instance in the image, obtaining a coarse segmentation result, employing a provided segmentation algorithm based on multi-scale feature fusion Deep Snake, and obtaining a final segmentation result; and calculating to obtain a contour vertex coordinate offset pointing to the boundary of the real instance, and adjusting the contour according to the offset to approach the contour boundary of the real portrait instance to obtain the portrait contour in the tourist image. According to the multi-scale feature fusion Deep Snake module, through improvement in two aspects of receptive field expansion and shallow feature enhancement, the defects of an original Deep Snake module are effectively overcome, and the accuracy and robustness of instance segmentation are improved, so that accurate statistics of tourists is realized, and automatic statistics of tourists can also be realized in scenic spots without ticket buying and ticket checking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention mainly relates to the fields of multi-scale feature fusion deep snake, image segmentation and automatic tourist number counting, and specifically relates to a tourist number counting method and system based on multi-scale feature fusion deep snake. Technical Background

[0002] Tourist number statistics are of great significance for understanding tourism market trends, formulating tourism policies, optimizing tourism resource allocation, etc., and play an important role in realizing intelligent tourism management.

[0003] Traditional tourism statistics methods often rely on ticket purchase / counting, manual surveys and data aggregation, which are time-consuming, labor-intensive, and have large errors. Especially for some scenic spots that do not require ticket purchase / ticket checking, tourist number statistics are more time-consuming and labor-intensive and difficult to achieve. With the development of artificial intelligence technology, the combination of artificial intelligence and tourism is one of the current research hotspots. Computer vision has become a technology that has attracted much attention in the field of artificial intelligence. It can convert digital images or videos into a form that can be recognized by computers, thereby helping machines to complete automated image recognition, classification, detection, tracking, segmentation and other tasks. According to the flow of people images taken by cameras in scenic spots, the number of tourists can be automatically counted by segmenting and detecting pedestrians in the image, so as to improve the accuracy and efficiency of tourist number statistics. The Deep Snake algorithm is a deep learning algorithm for image segmentation and contour detection. It can achieve efficient segmentation and detection while maintaining the fine structure of the contour. However, individuals in the crowd may present different sizes and shapes due to factors such as distance and perspective, so it is necessary to use multi-scale feature extraction technology to capture these changes. Through multi-scale feature extraction, the algorithm can capture portrait features at different scales. The proposed multi-scale feature fusion Deep Snake module effectively overcomes the shortcomings of the original Deep Snake module by expanding the receptive field and enhancing shallow features, and improves the accuracy and robustness of instance segmentation, thereby achieving accurate tourist number counting and automatic tourist number counting in scenic spots without ticket purchase and ticket inspection. Summary of the invention

[0004] In order to solve the above modeling defects, the present invention discloses a tourist number counting method and system based on multi-scale feature fusion deep snake.

[0005] To achieve the above purpose, the technical solution of the present invention is as follows:

[0006] A tourist number counting method based on multi-scale feature fusion deep snake includes the following steps:

[0007] Step 1: Train to obtain the multi-scale feature fusion Deep Snake module trained for the first time and the multi-scale feature fusion Deep Snake module trained for the second time;

[0008] Step 2: Obtain a tourist image containing multiple people, then use the CenterNet target detection network to detect the portrait instance in the tourist image, obtain the head rectangular detection frame, extract the midpoints of the four sides of the head rectangular detection frame, and connect them to form a diamond outline;

[0009] Step 3, uniformly sample S points on the diamond contour, use the multi-scale feature fusion Deep Snake module trained for the first time to calculate the coordinate offsets required for the four vertices to obtain the local extreme points of the avatar in the image, and obtain four extreme points. Use the four extreme points as the origin of the coordinate axis. The top extreme point draws a line segment with a length of 1 / 3 of the horizontal side length of the rectangular detection frame in the x-axis direction with itself as the center, and the bottom extreme point draws a line segment with a length of 1 / 3 of the horizontal side length of the rectangular detection frame in the x-axis direction with itself as the center. The leftmost extreme point draws a line segment with a length of 1 / 3 of the vertical side length of the rectangular detection frame in the y-axis direction with itself as the center, and the rightmost extreme point draws a line segment with a length of 1 / 3 of the vertical side length of the rectangular detection frame in the y-axis direction with itself as the center; connect the endpoints of the four line segments to form an octagon as the initial contour;

[0010] Step 4: uniformly sample m points on the initial contour to form an instance contour, and input the contour composed of the m points into the multi-scale feature fusion Deep Snake module trained for the second time to predict the offset that needs to be adjusted; adjust the instance contour by the contour vertex coordinate offset to approximate the real instance contour boundary, obtain the avatar contour in the tourist image, and then use the number of avatar contours in the tourist image as the number of tourists in the tourist image.

[0011] Further improvement, S=40, m=128.

[0012] As a further improvement, the multi-scale feature fusion Deep Snake module includes a backbone network, a multi-scale feature fusion network and a prediction network; the output of the backbone network is the input of the multi-scale feature fusion network, and the output of the multi-scale feature fusion network is the input of the prediction network; the prediction network outputs the avatar outline in the tourist image; the backbone network uses 8 layers of expanded circular convolution with convolution kernels of different scales to expand the receptive field and correct large prediction errors, and the 8 circular convolution layers are connected using jump connections.

[0013] As a further improvement, the 8-layer cyclic convolution data processing method of the backbone network is as follows:

[0014] 1.1) The first and second layers are convolution kernels with a circular convolution length of 1. The output of the first layer of circular convolution is:

[0015] f1=ReLU(Bn(circonv(x,n=1)))

[0016] Where ReLU represents the ReLU activation function, Bn represents the batch normalization operation, circonv represents the circular convolution operation function, n represents the convolution kernel length, x = [x1, x2, x3, ..., x 40 ] is 40 points uniformly sampled on the diamond contour, f1 is the output of the first layer of circular convolution; n is the length of circular convolution;

[0017] 1.2) The output f2 of the second layer of circular convolution is:

[0018] f2=f1+ReLU(Bn(circonv(f1,n=1)))

[0019] 1.3) The third and fourth layers are convolution kernels with a circular convolution length of 3, and the output f3 of the third layer of circular convolution is:

[0020] f3=f2+ReLU(Bn(circonv(f2,n=3)))

[0021] 1.4) The output of the 4th layer of circular convolution f4 combines the outputs of the 2nd and 3rd layers, and the calculation formula is:

[0022] f4=f2+f3+ReLU(Bn(circonv(f3,n=3)))

[0023] 1.5): The 5th and 6th layers are convolution kernels with a length of 4. The output of the 5th layer of circular convolution combines the outputs of the 2nd, 3rd and 4th layers. The calculation formula is

[0024] f5=f2+f3+f4+ReLU(Bn(circonv(f4,n=4)))

[0025] 1.6) The output of the 6th layer of circular convolution combines the outputs of the 2nd, 3rd, 4th and 5th layers, and the calculation formula is

[0026] f6=f2+f3+f4+f5+ReLU(Bn(circonv(f5,n=4)))

[0027] 1.7) The 7th and 8th layers are convolution kernels with a length of 6. The 7th layer circular convolution output f7 combines the outputs of the 2nd, 3rd, 4th, 5th and 6th layers. The calculation formula is

[0028] f7=f2+f3+f4+f5+f6+ReLU(Bn(circonv(f6,n=6)))

[0029] 1.8) The 8th layer circular convolution output f8 combines the outputs of the 2nd, 3rd, 4th, 5th, 6th and 7th layers, and the calculation formula is

[0030] f8=f2+f3+f4+f5+f6+f7+ReLU(Bn(circonv(f7,n=6)))

[0031] 1.9): The multi-scale feature fusion network splices all the features from the 2nd to the 8th layer in the backbone network together, and the calculation formula is

[0032] f Concat =concat(f2,f3,f4,f5,f6,f7,f8)

[0033] Among them, concat is a concatenation operation; f Concat The features after splicing.

[0034] As a further improvement, the data processing method of the multi-scale feature fusion network is as follows:

[0035] The concatenated features are processed through a 1*1 convolution layer and a maximum pooling layer, and then concatenated with the concatenated features. The calculation formula is:

[0036] f fusion =f Concat +maxpool(conv1d(f Concat ))

[0037] Among them, maxpool is the maximum pooling operation, conv1d is the 1*1 convolution operation, and f fusion It is the output of the feature fusion network.

[0038] As a further improvement, the prediction network uses three 1*1 convolutional layers to obtain the coordinate offset, and the calculation formula is:

[0039] f p1 =ReLU(conv1d(f fusion ))

[0040] f p2 =ReLU(conv1d(f p1 ))

[0041] f p3 =conv1d(f p2 )

[0042] Among them, f p1is the output of the first convolutional layer in the prediction network; f p2 is the output of the second convolutional layer in the prediction network; f p3 It is the output of the third convolutional layer in the prediction network, that is, the output of the multi-scale feature fusion Deep Snake module.

[0043] For further improvement, the multi-scale feature fusion Deep Snake module trained for the first time adopts a loss function L based on smooth l1 loss during training. df1

[0044]

[0045] Where N represents the number of sampled instance boundary contour vertices, are the coordinates of S points uniformly sampled on the diamond contour in step 2, represents the vertex coordinates of the label contour; smooth l1 represents the smoothhl1 loss, where the calculation formula is

[0046]

[0047] As a further improvement, the multi-scale feature fusion Deep Snake module trained for the second time adopts the following function loss as the loss function:

[0048]

[0049] in, are the coordinates of m points uniformly sampled on the initial contour in step 3.

[0050] S32, outputting the contour vertex coordinate offset obtained in step S31 as the final segmentation result to obtain the portrait segmentation result.

[0051] The advantages of the present invention are as follows:

[0052] The multi-scale feature fusion Deep Snake module of the present invention effectively overcomes the shortcomings of the original Deep Snake module by expanding the receptive field and enhancing shallow features, and improves the accuracy and robustness of instance segmentation, thereby achieving accurate tourist number counting and automatic tourist number counting in scenic spots without ticket purchase and ticket checking. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is a flow chart for tourist number statistics;

[0054] Figure 2 It is the portrait segmentation and detection process;

[0055] Figure 3It is a multi-scale feature fusion Deep Snake module. DETAILED DESCRIPTION

[0056] In order to facilitate the understanding of the present invention, the device of the present invention will be described more fully below with reference to the relevant drawings. Embodiments of the device are given in the drawings. However, the device can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.

[0057] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "principle", "method" and "theorem" should be understood in a broad sense, for example, it can be the equivalent principle, the superposition method, or Hooke's theorem, etc. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0058] The method and system for counting the number of tourists based on multi-scale feature fusion deep snake of the present invention mainly include: a camera taking a flow image of people (such as Figure 1 S11 in ), using a multi-scale feature fusion deep snake algorithm for portrait segmentation (such as Figure 1 S12 in), according to the segmentation results, mark the number of people (such as Figure 1 S13 in the figure), segmenting the newly appeared portrait in the picture (such as Figure 1 S14 in) and increase the number of marked people, and display the total number of people counted in real time (such as Figure 1 The system hardware includes a camera, a transmission cable, and a computer equipped with a segmentation algorithm. The result of the headcount is displayed in real time through a human-computer interaction interface.

[0059] The specific implementation includes the following steps:

[0060] (i) Take images of tourists with a camera installed at the exit of the scenic spot, use the CenterNet object detection network to detect portrait instances in the image, and extract the midpoints of the four sides of the detection frame to form a diamond outline;

[0061] (ii) 40 points are uniformly sampled on the diamond contour, and the multi-scale feature fusion Deep Snake module is used to find local extreme points. Then, the four extreme points with the smallest extreme values ​​are extended and connected in the horizontal and vertical directions to form an octagon as the initial contour;

[0062] (III) 128 points are uniformly sampled on the octagonal contour to ensure that it can retain most instance shapes. The octagonal contour composed of the sampling points is used as the starting point, and the multi-scale feature fusion Deep Snake module is used to predict the coordinate offset of the contour vertex pointing to the boundary of the real instance, and the contour is adjusted accordingly to approach the boundary of the real instance contour, and the portrait contour in the tourist image is obtained. The number of portrait segmentations in the image is counted to obtain the number of tourists.

[0063] The steps of using multi-scale feature fusion deep snake algorithm for portrait recognition are as follows Figure 2 As shown:

[0064] S21: Based on the tourist images captured by the camera, the images are transmitted in real time to a people counting system on a computer via a cable;

[0065] S22: Detect the portrait instance in the image through the CenterNet object detection network, extract the midpoints of the four sides of the detection box, and form a diamond outline;

[0066] S23: Sample 40 points uniformly on the diamond contour and use the multi-scale feature fusion Deep Snake module to find local extreme points. Figure 3 As shown in Figure 1, it includes a backbone network, a multi-scale feature fusion network, and a prediction network. The backbone network uses 8 layers of dilated circular convolution with convolution kernels of different scales to expand the receptive field and correct large prediction errors. The 8 circular convolution layers are connected using skip connections to retain shallow feature information that is passed to the deep network, so that the deformation module can better handle the segmentation task of small-scale instances and improve the overall segmentation performance. Figure 3 As mentioned above, in the backbone network with 8 layers of circular convolution, the first and second layers have convolution kernels with a circular convolution length of 1, and the output of the first layer of circular convolution is

[0067] f1=ReLU(Bn(circonv(x,n=1)))

[0068] Where ReLU represents the ReLU activation function, Bn represents the batch normalization operation, circonv represents the circular convolution operation function, n represents the convolution kernel length, x = [x1, x2, x3, ..., x 40 ] is 40 points uniformly sampled on the diamond contour, and f1 is the output of the first layer of circular convolution.

[0069] The output of the second layer of circular convolution is

[0070] f2=f1+ReLU(Bn(circonv(f1,n=1)))

[0071] S24: The 3rd and 4th layers are circular convolution kernels with a length of 3. The output of the 3rd layer circular convolution is

[0072] f3=f2+ReLU(Bn(circonv(f2,n=3)))

[0073] The output of the 4th layer of circular convolution combines the outputs of the 2nd and 3rd layers, and the calculation formula is:

[0074] f4=f2+f3+ReLU(Bn(circonv(f3,n=3)))

[0075] S25: The 5th and 6th layers are convolution kernels with a length of 4. The output of the 5th layer of circular convolution combines the outputs of the 2nd, 3rd and 4th layers. The calculation formula is

[0076] f5=f2+f3+f4+ReLU(Bn(circonv(f4,n=4)))

[0077] The output of the 6th layer of circular convolution combines the outputs of the 2nd, 3rd, 4th and 5th layers, and the calculation formula is:

[0078] f6=f2+f3+f4+f5+ReLU(Bn(circonv(f5,n=4)))

[0079] S26: The 7th and 8th layers are convolution kernels with a length of 6. The output of the 7th layer of circular convolution combines the outputs of the 2nd, 3rd, 4th, 5th and 6th layers. The calculation formula is:

[0080] f7=f2+f3+f4+f5+f6+ReLU(Bn(circonv(f6,n=6)))

[0081] The output of the 8th layer of circular convolution combines the outputs of the 2nd, 3rd, 4th, 5th, 6th and 7th layers, and the calculation formula is

[0082] f8=f2+f3+f4+f5+f6+f7+ReLU(Bn(circonv(f7,n=6)))

[0083] S27: The multi-scale feature fusion network concatenates all the features from the 2nd to the 8th layer in the backbone network. The calculation formula is:

[0084] f Concat =concat(f2,f3,f4,f5,f6,f7,f8)

[0085] Among them, concat is a concatenation operation.

[0086] After the features are concatenated, they are passed through a 1*1 convolution layer and a maximum pooling layer, and then concatenated with the concatenated features again. The calculation formula is

[0087] f fusion =f Concat +maxpool(conv1d(f Concat ))

[0088] Among them, maxpool is the maximum pooling operation, conv1d is the 1*1 convolution operation, and f fusion It is the output of the feature fusion network.

[0089] S28: The prediction network uses three 1*1 convolutional layers to obtain the coordinate offset, and the calculation formula is:

[0090] f p1 =ReLU(conv1d(f fusion ))

[0091] f p2 =ReLU(conv1d(f p1 ))

[0092] f p3 =conv1d(f p2 )

[0093] Among them, f p1 is the output of the first convolutional layer in the prediction network; f p2 is the output of the second convolutional layer in the prediction network; f p3 It is the output of the third convolutional layer in the prediction network, that is, the output of the multi-scale feature fusion Deep Snake module.

[0094] During the iterative training of the multi-scale feature fusion Deep Snake module, the loss function uses the loss function L based on the smooth l1 loss. df1

[0095]

[0096] Where N represents the number of sampled instance boundary contour vertices, are the coordinates of the contour vertices after the first segmentation, represents the vertex coordinates of the label contour. Smooth l1 represents the smooth l1 loss, where the calculation formula is

[0097]

[0098] S29: Extend and connect the four extreme points with the smallest extreme values ​​in the horizontal and vertical directions to form an octagon as the initial contour;

[0099] S30: 128 points are uniformly sampled on the octagonal outline to ensure that it can preserve most instance shapes.

[0100] S31: Using the octagonal contour composed of the sampling points as the starting point, the multi-scale feature fusion Deep Snake module is used to predict the coordinate offset of the contour vertex pointing to the boundary of the real instance, and the contour is adjusted accordingly to approach the boundary of the real instance contour. The following loss function is used in the training of the multi-scale feature fusion Deep Snake module:

[0101]

[0102] in, are the coordinates of m points uniformly sampled on the initial contour in step 3.

[0103] S32, outputting the contour vertex coordinate offset obtained in step S31 as the final segmentation result to obtain the portrait segmentation result.

[0104] A comparative experiment was conducted between the segmentation method based on multi-scale feature fusion Deep Snake and the traditional Deep Snake segmentation method. The overlap rate of the "predicted segmentation box" and the "real box" was calculated using loU, that is, the ratio of their intersection and union, AP vol The value is the average value of the average precision after the IoU threshold is divided into nine equal parts from 0.1 to 0.9, AP 50 It refers to the average precision when the IoU threshold is set to 0.5, AP 70 It refers to the average precision when the IoU threshold is set to 0.7. vol Value, AP 50 and AP 70 These three indicators are compared, and the comparison results are shown in Table 1. After expanding the receptive field and enhancing shallow features, the segmentation method based on multi-scale feature fusion Deep Snake has improved its accuracy by 1.6%.

[0105] Table 1 Comparison of segmentation accuracy between the multi-scale feature fusion deep snake segmentation method and the traditional deep snake segmentation method

[0106]

[0107] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention still falls within the scope of the technical solution of the present invention.

Claims

1. A tourist number counting method based on multi-scale feature fusion deep snake, characterized in that: The following steps are involved: Step 1: Train to obtain the multi-scale feature fusion Deep Snake module trained for the first time and the multi-scale feature fusion Deep Snake module trained for the second time; Step 2: Obtain a tourist image containing multiple people, then use the CenterNet target detection network to detect the portrait instance in the tourist image, obtain the head rectangular detection frame, extract the midpoints of the four sides of the head rectangular detection frame, and connect them to form a diamond outline; Step 3, uniformly sample S points on the diamond contour, use the multi-scale feature fusion DeepSnake module trained for the first time to calculate the coordinate offsets required for the four vertices to obtain the local extreme points of the avatar in the image, and obtain four extreme points. Use the four extreme points as the origin of the coordinate axis. The top extreme point draws a line segment with a length of 1 / 3 of the horizontal side length of the rectangular detection frame in the x-axis direction with itself as the center, and the bottom extreme point draws a line segment with a length of 1 / 3 of the horizontal side length of the rectangular detection frame in the x-axis direction with itself as the center. The leftmost extreme point draws a line segment with a length of 1 / 3 of the vertical side length of the rectangular detection frame in the y-axis direction with itself as the center, and the rightmost extreme point draws a line segment with a length of 1 / 3 of the vertical side length of the rectangular detection frame in the y-axis direction with itself as the center; connect the endpoints of the four line segments to form an octagon as the initial contour; Step 4: uniformly sample m points on the initial contour to form an instance contour, and input the contour composed of the m points into the multi-scale feature fusion Deep Snake module trained for the second time to predict the offset that needs to be adjusted; adjust the instance contour by the contour vertex coordinate offset to approximate the real instance contour boundary, obtain the avatar contour in the tourist image, and then use the number of avatar contours in the tourist image as the number of tourists in the tourist image.

2. The tourist number counting method based on multi-scale feature fusion deep snake as claimed in claim 1 is characterized in that: S=40, m=128.

3. The tourist number counting method based on multi-scale feature fusion deep snake as claimed in claim 2 is characterized in that: The multi-scale feature fusion Deep Snake module includes a backbone network, a multi-scale feature fusion network and a prediction network; the output of the backbone network is the input of the multi-scale feature fusion network, and the output of the multi-scale feature fusion network is the input of the prediction network; the prediction network outputs the head portrait outline in the tourist image; the backbone network uses 8 layers of expanded circular convolution with convolution kernels of different scales to expand the receptive field and correct large prediction errors, and the 8 circular convolution layers are connected using jump connections.

4. The tourist number counting method based on multi-scale feature fusion deep snake as claimed in claim 3 is characterized in that: The 8-layer circular convolution data processing method of the backbone network is as follows: 1.1) The first and second layers are convolution kernels with a circular convolution length of 1. The output of the first layer of circular convolution is: f1=ReLU(Bn(circonv(x,n=1))) Where ReLU represents the ReLU activation function, Bn represents the batch normalization operation, circonv represents the circular convolution operation function, n represents the convolution kernel length, x = [x1, x2, x3, ..., x 40 ] is 40 points uniformly sampled on the diamond contour, f1 is the output of the first layer of circular convolution; n is the length of circular convolution; 1.2) The output f2 of the second layer of circular convolution is: f2=f1+ReLU(Bn(circonv(f1,n=1))) 1.3) The third and fourth layers are convolution kernels with a circular convolution length of 3, and the output f3 of the third layer of circular convolution is: f3=f2+ReLU(Bn(circonv(f2,n=3))) 1.4) The output of the 4th layer of circular convolution f4 combines the outputs of the 2nd and 3rd layers, and the calculation formula is: f4=f2+f3+ReLU(Bn(circonv(f3,n=3))) 1.5): The 5th and 6th layers are convolution kernels with a length of 4. The output of the 5th layer of circular convolution combines the outputs of the 2nd, 3rd and 4th layers. The calculation formula is f5=f2+f3+f4+ReLU(Bn(circonv(f4,n=4))) 1.6) The output of the 6th layer of circular convolution combines the outputs of the 2nd, 3rd, 4th and 5th layers, and the calculation formula is f6=f2+f3+f4+f5+ReLU(Bn(circonv(f5,n=4))) 1.7) The 7th and 8th layers are convolution kernels with a length of 6. The 7th layer circular convolution output f7 combines the outputs of the 2nd, 3rd, 4th, 5th and 6th layers. The calculation formula is f7=f2+f3+f4+f5+f6+ReLU(Bn(circonv(f6,n=6))) 1.8) The 8th layer circular convolution output f8 combines the outputs of the 2nd, 3rd, 4th, 5th, 6th and 7th layers, and the calculation formula is f8=f2+f3+f4+f5+f6+f7+ReLU(Bn(circonv(f7,n=6))) 1.9): The multi-scale feature fusion network splices all the features from the 2nd to the 8th layer in the backbone network together, and the calculation formula is f Concat =concat(f2,f3,f4,f5,f6,f7,f8) Among them, concat is a concatenation operation; f Concat The features after splicing.

5. The tourist number counting method based on multi-scale feature fusion deep snake as claimed in claim 4 is characterized in that: The data processing method of the multi-scale feature fusion network is as follows: The concatenated features are processed through a 1*1 convolution layer and a maximum pooling layer, and then concatenated with the concatenated features. The calculation formula is: f fusion =f Concat +maxpool(conv1d(f Concat )) Among them, maxpool is the maximum pooling operation, conv1d is the 1*1 convolution operation, and f fusion It is the output of the feature fusion network.

6. The tourist number counting method based on multi-scale feature fusion deep snake as claimed in claim 5 is characterized in that: The prediction network uses three 1*1 convolutional layers to obtain the coordinate offset, and the calculation formula is: f p1 =ReLU(conv1d(f fusion )) f p2 =ReLU(conv1d(f p1 )) f p3 =conv1d(f p2 ) Among them, f p1 is the output of the first convolutional layer in the prediction network; f p2 is the output of the second convolutional layer in the prediction network; f p3 It is the output of the third convolutional layer in the prediction network, that is, the output of the multi-scale feature fusion Deep Snake module.

7. The tourist number counting method based on multi-scale feature fusion deep snake as claimed in claim 1 is characterized in that: The multi-scale feature fusion Deep Snake module trained for the first time uses a loss function L based on the smooth l1 loss during training. df1 Where N represents the number of sampled instance boundary contour vertices, are the coordinates of S points uniformly sampled on the diamond contour in step 2, represents the vertex coordinates of the label contour; smooth l1 represents the smoothhl1 loss, where the calculation formula is 8. The tourist number counting method based on multi-scale feature fusion deep snake as claimed in claim 1 is characterized in that: The multi-scale feature fusion Deep Snake module trained for the second time uses the following function loss as the loss function: in, are the coordinates of m points uniformly sampled on the initial contour in step 3. S32, outputting the contour vertex coordinate offset obtained in step S31 as the final segmentation result to obtain the portrait segmentation result.

Citation Information

Patent Citations

  • Crowd counting model training method and device, crowd counting method and device and server

    CN111046747A

  • Target contour tracking method and system based on correlation filtering and Deep Snake

    CN113658224A

  • Intelligent classroom analysis and identification method and device based on artificial intelligence

    CN116758487A

  • Depth instance segmentation method based on contour intersection-to-union ratio loss

    CN117253041A

  • People counting method, apparatus, and device based on facial recognition, and storage medium

    WO2020207038A1