Anti-occlusion human head 3D positioning method and system based on instance segmentation
Through example segmentation technology and migration image repair technology, combined with inverse proportional model, the problem of large error in the three-dimensional positioning of the human body is solved, and high-precision three-dimensional positioning of the human head is achieved.
Patent Information
- Application Number
- CN202411707387.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-11-27
AI Technical Summary
In the three-dimensional positioning of the human body, the error is large due to the change in the area of the human body bounding box, and the blocked human head bounding box cannot reflect the real depth information.
The three-dimensional positioning method of anti-occlusion human head based on instance segmentation is adopted, and the precise size is obtained through the human head instance segmentation technology, and the migration image repair technology is used to occlude and segment it. The three-dimensional positioning is combined with the inverse proportional model of the distance between the square of the area occupied by the head in the image and the camera optical axis direction.
It effectively improves the accuracy of the three-dimensional positioning of the human head, overcomes the problem that the human head area cannot reflect the real depth information under occlusion, and balances the calculation accuracy and speed.
Smart Images

Figure CN119205924B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to an anti-occlusion human head three-dimensional positioning method and system based on instance segmentation. Background Art
[0002] The three-dimensional positioning technology of personnel based on the video surveillance system plays an important role in maintaining public safety, optimizing building energy-saving solutions, and scheduling crowds.
[0003] In pictures or videos taken by a camera, the human body is larger when it is near and smaller when it is far away. Based on the above phenomenon, Moon et al. used the area of the human body bounding box detected by the Mask R-CNN algorithm to achieve 3D positioning of the human body root node based on a single image. However, the human body is a non-rigid body and moves frequently. The change of human posture has a huge impact on the area of the human body bounding box. Therefore, the area of the human body bounding box in the picture is not strictly larger when it is near and smaller when it is far away, which leads to a large error in the 3D positioning of the human body root node achieved using the area of the human body bounding box.
[0004] In the human body, the head is approximately a sphere, and the movement of the head will not significantly affect the shape of the head, so the head is approximately a spherical rigid body. The research results of Lian et al. show that in the picture taken by the camera, the radius of the head is inversely proportional to its distance from the camera in the direction of the camera optical axis. Based on the above research results, Zhao et al. combined the head detection model to realize the three-dimensional positioning of the head based on a single camera. However, on the one hand, the bounding box of the head given by the detection model is difficult to accurately fit the edge of the head, and on the other hand, the bounding box of the partially occluded head cannot reflect the real depth information of the head. Therefore, the three-dimensional coordinates of the head obtained by the above method still have certain errors. Summary of the invention
[0005] In order to solve the above problems, the present invention proposes an anti-occlusion human head three-dimensional positioning method and system based on instance segmentation, which adopts the human head instance segmentation technology to obtain the precise size of the human head in the picture; and the migration image restoration technology realizes the human head instance segmentation after de-occlusion. In order to balance the calculation accuracy and calculation speed, the present invention first detects and segments in the low-resolution image, and then further refines and repairs in the high-resolution image. Based on the above human head instance segmentation results and the inverse proportional model of the square root of the area occupied by the human head in the image and the distance between the human head in the direction of the camera optical axis and the camera, the three-dimensional positioning of the human head is performed.
[0006] In order to achieve the above object, the present invention adopts the following technical solution:
[0007] In a first aspect, the present invention provides an anti-occlusion human head three-dimensional positioning method based on instance segmentation, comprising:
[0008] Obtain a high-resolution video frame and its corresponding low-resolution video frame;
[0009] Input the low-resolution video frame into the feature extraction network to obtain shallow features and deep features;
[0010] Inputting the deep features into a feature fusion detection network to obtain fusion features and low-resolution detection results; the low-resolution detection results include instances with occluded human heads and instances without occluded human heads;
[0011] Input shallow features and fused features into the feature decoding segmentation network to obtain low-resolution feature maps and low-resolution head segmentation masks;
[0012] Based on the high-resolution video frame, the unobstructed head instance, the low-resolution feature map and the low-resolution head segmentation mask, the high-resolution unobstructed head instance segmentation result is obtained;
[0013] Based on the high-resolution video frame, the occluded head instance and the low-resolution head segmentation mask, the high-resolution de-occluded head instance segmentation result is obtained;
[0014] The high-resolution unoccluded head instance segmentation result and the high-resolution de-occluded head instance segmentation result are input into the three-dimensional positioning module to obtain the three-dimensional coordinates of the head.
[0015] Preferably, the low-resolution video frame is input into the feature extraction network to obtain shallow features and deep features, specifically including: the feature extraction network includes 6 network blocks connected in series, each network block outputs a feature map; the 3 feature maps output by the first 3 network blocks are shallow features, namely the first, second and third shallow features, and the 3 feature maps output by the last 3 network blocks are deep features; wherein, the first 2 network blocks each include 1 Fourier convolution layer, and the last 4 network blocks each include 1 image block fusion layer and 2 visual mamba layers.
[0016] Preferably, the feature fusion detection network includes a fast spatial pyramid pooling block, a bidirectional feature fusion block and a detection head;
[0017] Among them, the fast spatial pyramid pooling block is used to further extract and fuse high-level image features based on the three deep features; the bidirectional feature fusion block is used to fuse features of different scales to obtain the first, second and third fused features corresponding to the deep features; the fused features are respectively sent to their respective detection heads, which include two branches composed of convolutions, one branch is used to output the coordinates of the head bounding box, and the other branch is used to output the head category, where the category includes unoccluded head instances and occluded head instances.
[0018] Preferably, the step of inputting the shallow features and the fused features into a feature decoding segmentation network to obtain a low-resolution feature map and a low-resolution head segmentation mask specifically includes:
[0019] The feature decoding segmentation network includes three network blocks connected in series and a segmentation head;
[0020] The first network block takes the first fusion feature as input, passes through an image expansion layer, and then concatenates it with the third shallow feature in the channel dimension and sends it to a visual Mamba layer to output a feature map. ; The second network block starts with As input, after the upsampling layer, it is concatenated with the second shallow layer features and sent to a Fourier convolution layer to output the feature map ; The third network block starts with As input, after the upsampling layer, it is concatenated with the first shallow feature layer and sent to a Fourier convolution layer to output a low-resolution feature map. ; Split the header to As input, after two Fourier convolution layers, the output is a low-resolution head segmentation mask .
[0021] Preferably, the method of obtaining a high-resolution unobstructed head instance segmentation result based on the high-resolution video frame, the unobstructed head instance, the low-resolution feature map and the low-resolution head segmentation mask specifically includes:
[0022] Upsampling the low-resolution head segmentation mask and the low-resolution feature map to obtain a high-resolution head segmentation mask and a high-resolution feature map;
[0023] According to the high-resolution video frame, the unobstructed head instance is restored to high resolution, and the bounding box coordinates of the unobstructed head instance at high resolution are obtained;
[0024] The high-resolution video frame, the high-resolution head segmentation mask and the high-resolution feature map are cropped using the bounding box coordinates of the unobstructed head instance at high resolution to obtain the high-resolution local image, local feature map and local segmentation mask corresponding to each human head instance;
[0025] The feature map obtained after deformable Fourier convolution of the high-resolution local picture corresponding to each head instance is spliced with the local feature map of the corresponding head instance and sent to the deformable Fourier convolution residual block. The output feature is superimposed on the local segmentation mask of the corresponding head instance to obtain a high-resolution unobstructed head instance segmentation result.
[0026] Preferably, the step of obtaining a high-resolution de-occluded head instance segmentation result based on the high-resolution video frame, the occluded head instance and the low-resolution head segmentation mask specifically includes:
[0027] Upsample the low-resolution head segmentation mask to obtain a high-resolution head segmentation mask;
[0028] According to the high-resolution video frame, the occluded head instance is restored to high resolution, and the coordinates of the bounding box of the occluded head instance at high resolution are obtained;
[0029] Expand the bounding box coordinates of the occluded head instance at high resolution, and use the expanded bounding box coordinates to crop the high-resolution video frame and the high-resolution head segmentation mask to obtain the high-resolution local image and local segmentation mask corresponding to each occluded head instance;
[0030] The high-resolution local image and local segmentation mask corresponding to each occluded head instance are concatenated and sent to the UNet segmentation network to output the occluder mask in the high-resolution local image corresponding to the occluded head instance;
[0031] The high-resolution local image corresponding to the occluded head instance is fused with the occlusion mask output in the previous step and input into the LaMa-based de-occlusion head segmentation network to obtain the high-resolution de-occluded head instance segmentation result.
[0032] Preferably, the high-resolution unobstructed head instance segmentation result and the high-resolution de-occluded head instance segmentation result are input into a three-dimensional positioning module to obtain the three-dimensional coordinates of the head, which specifically includes:
[0033] The area occupied by each head instance in the image is obtained from the high-resolution unobstructed head instance segmentation results and the high-resolution de-occluded head instance segmentation results, and the distance between the head and the camera in the direction of the camera optical axis is calculated using the inverse proportional model of the square root of the area occupied by the head in the image and the distance between the head and the camera in the direction of the camera optical axis;
[0034] The three-dimensional coordinates of the head in the camera coordinate system are obtained according to the distance between the head and the camera in the direction of the camera optical axis, the camera internal parameters and the bounding box coordinates of the head;
[0035] According to the obtained three-dimensional coordinates of the head in the camera coordinate system and the camera extrinsic parameters, the three-dimensional coordinates of the head in the world coordinate system are calculated, and the three-dimensional positioning result of the person in the single camera is obtained.
[0036] In a second aspect, the present invention provides an anti-occlusion human head three-dimensional positioning system based on instance segmentation, comprising:
[0037] An image acquisition module, used to acquire a high-resolution video frame and its corresponding low-resolution video frame;
[0038] A low-resolution feature extraction module is used to input low-resolution video frames into a feature extraction network to obtain shallow features and deep features;
[0039] A low-resolution detection module, used to input deep features into a feature fusion detection network to obtain fusion features and low-resolution detection results; the low-resolution detection results include instances with occluded human heads and instances without occluded human heads;
[0040] The low-resolution segmentation module is used to input shallow features and fused features into the feature decoding segmentation network to obtain a low-resolution feature map and a low-resolution head segmentation mask;
[0041] A high-resolution unobstructed segmentation refinement module is used to obtain a high-resolution unobstructed head instance segmentation result based on a high-resolution video frame, an unobstructed head instance, a low-resolution feature map, and a low-resolution head segmentation mask;
[0042] A high-resolution deocclusion restoration segmentation module is used to obtain a high-resolution deocclusion head instance segmentation result based on a high-resolution video frame, an occluded head instance, and a low-resolution head segmentation mask;
[0043] The three-dimensional positioning module is used to input the high-resolution unobstructed head instance segmentation result and the high-resolution de-occluded head instance segmentation result into the three-dimensional positioning module to obtain the three-dimensional coordinates of the head.
[0044] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for anti-occlusion human head three-dimensional positioning based on instance segmentation described in the first aspect.
[0045] In a fourth aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the method for anti-occluded three-dimensional human head positioning based on instance segmentation described in the first aspect are implemented.
[0046] Compared with the prior art, the present invention has the following beneficial effects:
[0047] The present invention adopts the head instance segmentation technology to obtain the precise size of the head in the picture, and then uses the square root of the area occupied by the head in the image and the inverse proportional model of the distance between the head in the direction of the camera optical axis and the camera to perform three-dimensional positioning of the head, effectively improving the three-dimensional positioning accuracy of the head based on a single camera; the de-occlusion segmentation technology of the migration image restoration is adopted to obtain the head instance segmentation result after de-occlusion, effectively overcoming the problem that the head area obtained by simply relying on the instance segmentation technology in the case of occlusion cannot reflect the real depth information of the head; after detection and segmentation in the low-resolution image, further refinement and restoration are performed in the high-resolution local image, effectively balancing the calculation accuracy and speed.
[0048] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their description are used to explain the present invention but do not constitute a limitation of the present invention.
[0050] Figure 1 A main flow chart of an anti-occlusion human head three-dimensional positioning method based on instance segmentation provided by an embodiment of the present invention;
[0051] Figure 2 An image acquisition module provided by an embodiment of the present invention;
[0052] Figure 3 A structural diagram of a low-resolution feature extraction module provided in an embodiment of the present invention;
[0053] Figure 4 A structural diagram of a Fourier convolution layer provided in an embodiment of the present invention;
[0054] Figure 5 A schematic diagram of the process of the image block fusion layer provided by an embodiment of the present invention;
[0055] Figure 6 A structural diagram of the visual mamba layer provided by an embodiment of the present invention;
[0056] Figure 7 A structural diagram of a low-resolution detection module provided in an embodiment of the present invention;
[0057] Figure 8 A structural diagram of a fast spatial pyramid pooling layer provided by an embodiment of the present invention;
[0058] Fig. 9 A structural diagram of a low-resolution segmentation module provided in an embodiment of the present invention;
[0059] Fig.10 A structural diagram of a high-resolution unobstructed segmentation and refinement module provided in an embodiment of the present invention;
[0060] Fig.11 A structural diagram of a deformable Fourier convolution layer provided in an embodiment of the present invention;
[0061] Fig.12 A structural diagram of a deformable Fourier convolution residual block provided by an embodiment of the present invention;
[0062] Fig.13 A structural diagram of a high-resolution de-occlusion and restoration segmentation module provided in an embodiment of the present invention;
[0063] Fig.14 The effect diagram of unobstructed three-dimensional positioning of a human head provided by an embodiment of the present invention;
[0064] Fig.15 This is a rendering of the three-dimensional positioning of an obstructed human head provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0065] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0066] Embodiment 1
[0067] like Figure 1 As shown, this embodiment discloses an anti-occlusion human head three-dimensional positioning method based on instance segmentation, comprising the following steps:
[0068] S1: Obtain a high-resolution video frame and its corresponding low-resolution video frame;
[0069] S2: Input the low-resolution video frame into the feature extraction network to obtain shallow features and deep features;
[0070] S3: Inputting the deep features into a feature fusion detection network to obtain fusion features and low-resolution detection results; the low-resolution detection results include instances with occluded human heads and instances without occluded human heads;
[0071] S4: Input shallow features and fused features into the feature decoding segmentation network to obtain low-resolution feature maps and low-resolution head segmentation masks;
[0072] S5: Based on the high-resolution video frame, the unobstructed head instance, the low-resolution feature map and the low-resolution head segmentation mask, a high-resolution unobstructed head instance segmentation result is obtained;
[0073] S6: Based on the high-resolution video frame, the occluded head instance and the low-resolution head segmentation mask, the high-resolution de-occluded head instance segmentation result is obtained;
[0074] S7: Input the high-resolution unobstructed head instance segmentation result and the high-resolution de-occluded head instance segmentation result into a three-dimensional positioning module to obtain the three-dimensional coordinates of the head.
[0075] Next, combine Figure 1 , a method for anti-occlusion human head three-dimensional positioning based on instance segmentation disclosed in this embodiment is described in detail.
[0076] In S1, specifically, a surveillance video is obtained from a surveillance camera.
[0077] Currently, thanks to the advancement of imaging technology and network transmission speed, surveillance cameras often output high-definition or ultra-high-definition video. For example, a surveillance camera worth only RMB 100 on the market can output 3840 2160 resolution picture. If the high-resolution video frame is directly sent to the image processing network, on the one hand, it will greatly increase the network calculation amount, and on the other hand, the image processing network needs to have higher global processing capabilities.
[0078] In order to balance the computational accuracy and computational speed, this embodiment first performs detection and segmentation in a low-resolution image, and then further performs refinement and repair in a high-resolution image. Therefore, this embodiment first obtains a high-resolution video frame from a single surveillance camera. , and then perform downsampling to obtain the corresponding low-resolution video frame , the process is as follows Figure 2 shown.
[0079] Used for preliminary detection and segmentation of human heads, For further refinement and repair. The height and width are and , The height and width are and Therefore, the downsampling rate in the vertical direction is , the downsampling rate in the horizontal direction is .
[0080] In S2, Figure 3 As shown in Figure 1, the feature extraction network is a deep neural network composed of a Fourier convolution layer, a visual Mamba layer, and an image block fusion layer.
[0081] The feature extraction network consists of 6 network blocks connected in series, among which the first and second network blocks each contain a Fourier convolution layer, and the third, fourth, fifth, and sixth network blocks each contain an image block fusion layer and two visual mamba layers.
[0082] The image feature maps output by the above 6 network blocks are , , , , , .in, is the first shallow feature, whose height and width are equal to and ; is the second shallow feature, whose height and width are equal to and ; is the third shallow feature, whose height and width are equal to and . , , They are all deep features. The height and width are equal to and , The height and width are equal to and , The height and width are equal to and .
[0083] As a specific implementation, the Fourier convolutional layer network structure is as follows: Figure 4 As shown in Figure 2, it can be seen that the Fourier convolution layer includes two feature processing channels. After the convolution process, the right channel The features after convolution are added and fused; the features of the right channel are processed by frequency domain transformation and combined with the features of the left channel by the second The features after convolution processing are added and fused.
[0084] The main idea of frequency domain transformation is to first perform Fourier transformation on the features, transform the spatial features to the frequency domain, then perform convolution on the frequency domain features, and finally send the convolved features to the inverse Fourier transform to transform the features back to the spatial domain. Since a slight change in the frequency domain can change the features of the entire spatial domain, the frequency domain transformation has the ability to extract global features.
[0085] like Figure 4 In the Fourier convolution layer shown, the left channel only involves the normal Convolution, so the left channel is used to extract local features, and the right channel involves ordinary While performing convolution, frequency domain transformation is also used, so the right channel is used to extract global features. The addition and fusion of the left and right channel features helps to fuse global features with local features. In this embodiment, Fourier convolution is used in the first and second network blocks mainly to increase the global feature perception ability of the network in the early stage.
[0086] As a specific implementation, the basic principle of the image block fusion layer is as follows: Figure 5 As shown in Figure 2, the image block fusion layer mainly includes three steps: splitting, splicing, and channel transformation. In the splitting step, first, the original feature map is divided into many local area, such as Figure 5 As shown, assuming that the upper left position of each local area is marked as 1, the upper right position is marked as 2, the lower left position is marked as 3, and the lower right position is marked as 4, the size of the original feature map is ; Then, combine all the positions marked as 1 into a new size of The feature maps of the other are similar, and a total of 4 sizes are obtained In the concatenation step, the above four feature maps are concatenated in the channel dimension to obtain a feature map of size In the channel transformation step, we use The convolution of the above size is The feature map is transformed into a size of In summary, the main function of the image block fusion layer is to halve the length and width of the feature map and double the number of channels.
[0087] As a specific implementation, the visual Mamba layer network structure is as follows: Figure 6 As shown in the figure, the visual mamba layer includes two residual structures: the first residual structure includes layer normalization and two-dimensional selective scanning processing; the second residual structure includes layer normalization and multi-layer perceptron. In the above structure, the two-dimensional selective scanning processing is the key to the visual mamba layer, which transforms the image feature map into a sequence of feature blocks, then uses selective scanning to process the above feature block sequence, and finally reshapes the processed feature block sequence into a feature map.
[0088] In S3, specifically, the feature fusion detection network is as follows Figure 7 As shown in Figure 1, it is a deep neural network composed of a fast spatial pyramid pooling block, a bidirectional feature fusion block, and a detection head. The network is based on the deep feature map output by the feature extraction module. , , As input, the three inputs are sent to the bidirectional feature fusion module after fast spatial pyramid pooling, and the fused feature map is output. , , ( , , The sizes are , , same); among which, is the first fusion feature, is the second fusion feature, It is the third fusion feature. , , They are sent to their respective detection heads respectively, and the low-resolution head detection results are output. ,in The first Normalized coordinates of the bounding box of a person's head, For the detected The categories of human heads are divided into two categories: covered and uncovered. is the total number of human heads detected.
[0089] As a specific implementation, the fast spatial pyramid pooling block network structure is as follows: Figure 8 As shown in Figure 2, it can be seen that the fast spatial pyramid pooling block is After the convolution, it is sent to 3 The maximum pooling layer of the convolution and pooling feature maps are concatenated in the channel dimension and finally The fast spatial pyramid pooling block is used to further extract and fuse higher-level image features. Figure 7 As shown on the left, the feature map , , The feature maps output after fast spatial pyramid pooling are , , .
[0090] As a specific implementation, the bidirectional feature fusion block network structure is as follows: Figure 7 As shown in the middle part of . It can be seen that the bidirectional feature fusion block includes the left upward fusion channel and the right downward fusion channel.
[0091] In the left upward fusion channel, the feature map After the visual Mamba layer, the feature map is obtained , The image block is sent to the expansion layer to expand the height and width of the feature map, and then combined with the feature map The spliced features are spliced in the channel dimension, and the spliced features are passed through the visual Mamba layer to obtain the feature map , The image block is sent to the expansion layer to expand the height and width of the feature map, and then combined with the feature map Concatenate on the channel dimension to get the feature map .
[0092] In the downward fusion channel on the right, the feature map After the visual Mamba layer, the first fusion feature map is obtained , The image block is sent to the fusion layer to compress the feature map height and width, and then combined with the feature map The spliced features are spliced in the channel dimension, and the spliced features are passed through the visual Mamba layer to obtain the second fusion feature map , The image block is sent to the fusion layer to compress the feature map height and width, and then combined with the feature map The spliced features are spliced in the channel dimension, and the spliced features are passed through the visual Mamba layer to obtain the third fusion feature map .
[0093] The upward fusion channel on the left mainly integrates high-level features into low-level features, and the downward fusion channel on the right mainly integrates low-level features into high-level features, thereby fully integrating features of different scales. The visual mamba layer and the image block fusion layer have been introduced in detail in the previous article and will not be repeated here. The image block expansion layer performs the inverse process of the image block fusion layer. Its main function is to double the length and width of the feature map and halve the number of channels.
[0094] As a specific implementation manner, the detection head is as follows Figure 7 Each detection head includes two branches consisting of convolutions, one branch is used to output the normalized coordinates of the head bounding box, and the other branch is used to output the head category.
[0095] In S4, specifically, the feature decoding segmentation network is as follows Fig. 9 As shown in the figure, the feature decoding segmentation network consists of an image block expansion layer, a visual mamba layer, an upsampling layer, and a Fourier convolution layer. The network consists of three network blocks and a segmentation head, where the first network block is based on the first fused feature map. As input, after one image expansion layer, it is combined with the third shallow feature Concatenate them in the channel dimension and send them to a visual Mamba layer to output the feature map ; The second network block starts with As input, after the upsampling layer, it is combined with the second shallow feature Splice them together and send them into a Fourier convolution layer to output the feature map ; The third network block starts with As input, after the upsampling layer, it is combined with the first shallow feature The concatenation is sent to a Fourier convolution layer to output a low-resolution feature map. ; Split the header to As input, after two Fourier convolution layers, the output is a low-resolution human head segmentation mask , The height and width are equal and A binary image, where the pixel 0 represents the background and the pixel 1 represents the foreground of the human head.
[0096] In S5, a high-resolution unoccluded head instance segmentation result is obtained based on the high-resolution video frame, the unoccluded head instance, the low-resolution feature map and the low-resolution head segmentation mask.
[0097] Specifically, Fig.10As shown. This module monitors video frames at high resolution , Human head segmentation mask at low resolution , feature map , the normalized coordinates of the low-resolution bounding box of the unobstructed head instance in the low-resolution head detection result As input, Indicates the number of heads that are classified as unobstructed in the low-resolution head detection results. Normalized coordinates of the bounding box of a person's head, It indicates the number of heads classified as unobstructed in the head detection results, and outputs the refined unobstructed head instance segmentation results at high resolution. Fig.10 The high-resolution unobstructed segmentation refinement module is shown in Take an unobstructed human head instance as an example. The specific steps are as follows:
[0098] Step S501: Segmentation mask of human head at low resolution and feature map Upsampling is performed, and the segmentation mask after upsampling is , the feature map is , and The height and width of high-resolution video frames Consistent, that is and ;
[0099] Step S502: high-resolution video frames Height and width , normalize the low-resolution bounding box coordinates of the unobstructed head instance in the low-resolution head detection result Restore to high resolution and obtain the bounding box coordinates of the unobstructed head instance at high resolution ,in Indicates high resolution The bounding box coordinates of unoccluded head instances, Indicates the number of heads classified as unobstructed in the head detection results;
[0100] Step S503: Utilize Cropping high-resolution video frames ,Feature map and segmentation mask , obtain the high-resolution local image corresponding to each human head instance , feature map and segmentation mask ,in, Indicates from The first High-resolution partial image of a human head. Indicates from The first Local feature map of the human head, Indicates from The first Individual head local segmentation mask;
[0101] Step S504: The feature map obtained after deformable Fourier convolution of the high-resolution local image corresponding to each head instance and the local feature map of the corresponding head instance are spliced and sent to the deformable Fourier convolution residual block, and the output feature is superimposed on the local segmentation mask of the corresponding head instance to complete the refinement of the segmentation mask of each unobstructed head instance at high resolution. The segmentation mask refinement result is used to further refine the bounding box coordinates of the corresponding head instance, and the refined unobstructed head instance segmentation result at high resolution is output. The refined unobstructed head instance segmentation result at high resolution includes the refined bounding box coordinates of each unobstructed head instance at high resolution and the refined segmentation mask at high resolution. Fig.10 First As an example, an unobstructed human head instance is shown. The feature map obtained after the deformable Fourier convolution and The concatenated features are then fed into two deformable Fourier convolution residual blocks, and the output features are superimposed to After high-resolution refinement, Segmentation mask of unobstructed head instances , Further used to adjust , after obtaining high-resolution refinement Bounding box coordinates of unobstructed head instances .
[0102] As a specific implementation method, the network structure of the deformable Fourier convolution is as follows: Fig.11 The instance segmentation refinement module requires the feature extraction network to be sensitive to the target shape. Therefore, this embodiment uses deformable convolution to replace the original Figure 4 The ordinary convolution in Fourier convolution is shown in Figure 1. The above deformable Fourier convolution residual block network structure is as follows: Fig.12 shown.
[0103] The output of the high-resolution unoccluded segmentation refinement module is the unoccluded head instance segmentation result after refinement at high resolution. ,in After high-resolution refinement The bounding box coordinates of unoccluded head instances, After high-resolution refinement Segmentation masks of unobstructed head instances, is the number of unobstructed head instances.
[0104] In S6, based on the high-resolution video frame, the occluded head instance and the low-resolution head segmentation mask, a high-resolution de-occluded head instance segmentation result is obtained.
[0105] Specifically, Fig.13 As shown. This module monitors video frames at high resolution , head segmentation mask at low resolution , the normalized coordinates of the low-resolution bounding box of the head instance classified as occluded in the low-resolution head detection result As input, Indicates the number of heads that are classified as occluded in the low-resolution head detection results. Normalized coordinates of the bounding box of a person's head, Indicates the number of heads classified as occluded in the head detection results, and outputs the head instance segmentation results after deocclusion with high resolution. The specific steps are as follows:
[0106] Step S601: Segmentation mask of the head at low resolution Upsampling is performed, and the segmentation mask after upsampling is , The height and width of high-resolution video frames Consistent, that is and ;
[0107] Step S602: high-resolution video frames Height and width , normalize the low-resolution bounding box coordinates of the occluded head instance in the low-resolution head detection result Restore to high resolution and obtain the bounding box coordinates of the occluded head instance at high resolution ;
[0108] Step S603: Expansion , and use the dilated bounding box coordinates to crop the high-resolution video frame and segmentation mask , respectively obtain the high-resolution local image corresponding to each occluded head instance and local segmentation mask ,in Indicates from The first A high-resolution partial image with an occluded head. Indicates from The first A local segmentation mask with an occluded head;
[0109] Step S604: The high-resolution local image and local segmentation mask corresponding to each occluded head instance are concatenated and sent to the UNet segmentation network to output the occlusion mask in the high-resolution local image corresponding to the occluded head instance. Fig.13 First Take an example with an occluded head as an example. and The occluder mask is spliced in the channel dimension and then fed into the UNet occluder segmentation network to output the occluder mask in the high-resolution local image corresponding to the occluded head instance. - ;
[0110] Step S605: The high-resolution local image corresponding to the occluded head instance is fused with the occlusion mask output in the previous step, and input into the LaMa-based de-occlusion head segmentation network, and the high-resolution de-occlusion head instance segmentation result is output. The segmentation result further adjusts the bounding box coordinates of the corresponding head instance, and the high-resolution de-occlusion head instance segmentation result is output. The high-resolution de-occlusion head instance segmentation result includes the high-resolution bounding box coordinates and high-resolution segmentation mask of each occluded head instance after de-occlusion. Fig.13 First Take an example of an occluded head as an example, the occluder mask output by the UNet network and Fusion (fusion method is and The multiplied image is the same as splicing), and then input the LaMa-based head segmentation network to output the head segmentation result after deocclusion , Further used to adjust , get the first Bounding box coordinates of occluded head instances .
[0111] The LaMa-based head segmentation network is migrated from the LaMa image restoration network proposed by Suvorov et al. In this embodiment, the LaMa-based head segmentation network adopts the same network structure as the LaMa image restoration network, and uses the parameters of the LaMa image restoration model trained under the face data CelebA-HQ to initialize the network during training. The output of the network is changed from the restoration image to the head segmentation mask, and retraining is performed.
[0112] In S7, as a specific implementation, the output of the de-occluded head instance segmentation module is a high-resolution de-occluded head instance segmentation result. ,in For high resolution The bounding box coordinates of the deoccluded head instance, For high resolution The segmentation mask of the head instance after deocclusion, is the number of occluded head instances.
[0113] As a specific implementation method, specifically, the three-dimensional positioning module outputs the three-dimensional coordinates of the center of the human head according to the internal and external parameter matrix of the camera calibration, as well as the unobstructed human head instance segmentation results after refinement at high resolution and the human head instance segmentation results after deocclusion at high resolution, thereby completing the three-dimensional positioning of the human head.
[0114] The camera intrinsic parameters include the product of the scale factor between the pixels in the image and the actual physical size and the focal length of the camera. , the horizontal coordinate of the camera optical center in the image , the vertical coordinate of the camera optical center in the image ; Camera external parameters include the rotation matrix between the camera coordinate system and the world coordinate system , the translation vector between the camera coordinate system and the world coordinate system ; The size is , The size is .
[0115] In step S7, the specific steps of three-dimensional positioning are:
[0116] Step S701: obtaining the area occupied by each head instance in the image from the high-resolution unobstructed head instance segmentation result and the high-resolution de-occluded head instance segmentation result, and calculating the distance between the head and the camera in the direction of the camera optical axis by using the square root of the area occupied by the head in the image and the inverse proportional model of the distance between the head and the camera in the direction of the camera optical axis;
[0117] Step S702: Obtain the three-dimensional coordinates of the head in the camera coordinate system according to the distance between the head and the camera in the direction of the camera optical axis, the camera internal parameters and the bounding box coordinates of the head;
[0118] Step S703: Calculate the three-dimensional coordinates of the head in the world coordinate system based on the obtained three-dimensional coordinates of the head in the camera coordinate system and the camera external parameters, so as to obtain the three-dimensional positioning result of the person in the single-camera environment.
[0119] In the study of Zhao et al., it has been proved that the distance between the human head and the camera in the direction of the camera optical axis is inversely proportional to the height of the human head in the image as shown below:
[0120] (1)
[0121] in, The product of the scale factor between the pixels in the image and the actual physical size and the focal length of the camera, Indicates the distance between the human head and the camera in the direction of the camera optical axis. Indicates the height of the human head in the image, Indicates the actual physical height of the human head. Based on the above research results, Zhao et al. combined the head detection model to achieve three-dimensional positioning of the head based on a single camera. However, the head boundary box given by the detection model is difficult to accurately fit the head boundary. Therefore, the three-dimensional coordinates of the head obtained according to the above method still have certain errors. In order to overcome the above problems, this embodiment uses the high-resolution head instance segmentation result that more accurately fits the head boundary to describe the size of the head in the image.
[0122] Assuming that the human head is a sphere, the height of the human head is close to the diameter of the human head, and the area of the human head is proportional to the square of the diameter of the human head. Based on this and formula (1), it can be deduced that the distance between the human head and the camera in the direction of the camera optical axis is approximately inversely proportional to the square root of the area of the human head in the image, which can be expressed as:
[0123] (2)
[0124] in, Indicates the area of the human head image in the image; represents the real cross-sectional area of the human head. In this embodiment, Obtained through camera calibration, The unobstructed head instance segmentation result after refinement at high resolution and the head instance segmentation result after deocclusion at high resolution are obtained. , is the actual physical diameter of the human head. It is equal to the height of the person being tested divided by 7.
[0125] For the convenience of description, the unobstructed head instance segmentation results after refinement at high resolution and the head instance segmentation results after deocclusion at high resolution are combined into one set, namely, the head instance segmentation results at high resolution ,in represents the total number of head instances at high resolution, Indicates high resolution Segmentation mask of a head instance, Indicates high resolution The bounding box coordinates of the individual head instances, , For the The horizontal coordinate of the upper left corner of the bounding box of the individual head instance, For the The ordinate of the upper left corner of the bounding box of the head instance. For the The width of the bounding box of the individual head instances, For the The height of the bounding box of a head instance. The distance between the head instance and the camera along the camera optical axis The calculation formula is as follows:
[0126] (3)
[0127] in, Indicates high resolution The area of a human head instance.
[0128] In step S702, the coordinates of the head in the camera coordinate system include X-axis coordinates, Y-axis coordinates and Z-axis coordinates, where the Z-axis is consistent with the optical axis direction of the camera, and the Z-axis coordinates of the head in the camera coordinate system are equal to the distance between the head and the camera in the direction of the camera optical axis, and the plane formed by the X-axis and the Y-axis is perpendicular to the optical axis of the camera. The 3D coordinates of the head instance in the camera coordinate system For example, according to the camera imaging principle, the calculation formula of the three-dimensional coordinates is as follows:
[0129] (4)
[0130] Among them, the horizontal coordinate of the camera optical center in the image is , the vertical coordinate of the camera optical center in the image and Obtained through camera calibration, , , , , Provided by the high-resolution head instance segmentation results, , is the actual physical diameter of the human head. It is equal to the height of the person being tested divided by 7.
[0131] In step S703, calculate the The coordinates of the head of the human body in the world coordinate system , the calculation formula is as follows:
[0132] (5)
[0133] in, is the three-dimensional coordinate of the head in the camera coordinate system calculated in step S702; is the rotation matrix between the world coordinate system and the camera coordinate system, with size , is the translation vector between the world coordinate system and the camera coordinate system, with a size of , and It belongs to the camera external parameter and is obtained through camera calibration. Therefore, the three-dimensional coordinates of the human head in the world coordinate system can be calculated by the above formula.
[0134] The 3D positioning effect of an unobstructed head is as follows Fig.14 As shown. As can be seen in the figure, the head instance segmentation mask at low resolution is very rough, and the positioning of the head contour is not accurate enough. After segmentation and refinement at high resolution, the contour of the head can be accurately captured, and the segmented head boundary is more accurate. According to the head area and the coordinates of the center point of the head in the instance segmentation mask after high-resolution refinement, the three-dimensional coordinates of the head instance in the world coordinate system can be further calculated to be (1.26, 1.50, 2.41) (unit: meter). The true value of the three-dimensional coordinates of the head in the figure is (1.2, 1.58, 2.4), and the error with the three-dimensional coordinates calculated by the method shown in this embodiment is (0.06, -0.08, 0.01). It can be seen that the method provided in this embodiment can achieve accurate three-dimensional positioning of unobstructed heads.
[0135] The 3D positioning effect with occluded head is as follows Fig.15 As shown. As can be seen from the figure, when the head is partially occluded, the segmentation module cannot effectively identify the occluded part, resulting in the segmentation mask area being unable to accurately reflect the distance from the head to the camera. After de-occlusion and repair segmentation at high resolution, the occlusion can be accurately removed and the complete head mask can be repaired. According to the head area and the coordinates of the center point of the head in the high-resolution de-occlusion and repair segmentation mask, the three-dimensional coordinates of the head instance in the world coordinate system can be further calculated as (-0.63, 1.83, 2.52) (unit: meter). The real three-dimensional coordinates of the head in the figure are (-0.6, 1.58, 2.4), and the error with the three-dimensional coordinates calculated by the method shown in this embodiment is (-0.03, 0.25, 0.12). It can be seen that the method provided in this embodiment can more accurately achieve three-dimensional positioning of the occluded head.
[0136] This specific embodiment uses instance segmentation technology to obtain the precise size of the head in the image. In order to balance the calculation accuracy and calculation speed, detection and segmentation are first performed in the low-resolution image, and then further refinement and repair are performed in the high-resolution image. On the other hand, the image restoration technology is migrated to obtain the head instance segmentation result after de-occlusion, and then based on the above head instance segmentation result and the inverse proportional model of the square root of the area occupied by the head in the image and the distance between the head and the camera in the direction of the camera optical axis, high-precision three-dimensional positioning of the head is performed.
[0137] Embodiment 2
[0138] This embodiment provides an anti-occlusion human head three-dimensional positioning system based on instance segmentation, including:
[0139] An image acquisition module, used to acquire a high-resolution video frame and its corresponding low-resolution video frame;
[0140] A low-resolution feature extraction module is used to input low-resolution video frames into a feature extraction network to obtain shallow features and deep features;
[0141] A low-resolution detection module, used to input deep features into a feature fusion detection network to obtain fusion features and low-resolution detection results; the low-resolution detection results include instances with occluded human heads and instances without occluded human heads;
[0142] The low-resolution segmentation module is used to input shallow features and fused features into the feature decoding segmentation network to obtain a low-resolution feature map and a low-resolution head segmentation mask;
[0143] A high-resolution unobstructed segmentation refinement module is used to obtain a high-resolution unobstructed head instance segmentation result based on a high-resolution video frame, an unobstructed head instance, a low-resolution feature map, and a low-resolution head segmentation mask;
[0144] A high-resolution deocclusion restoration segmentation module is used to obtain a high-resolution deocclusion head instance segmentation result based on a high-resolution video frame, an occluded head instance, and a low-resolution head segmentation mask;
[0145] The three-dimensional positioning module is used to input the high-resolution unobstructed head instance segmentation result and the high-resolution de-occluded head instance segmentation result into the three-dimensional positioning module to obtain the three-dimensional coordinates of the head.
[0146] Embodiment 3
[0147] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the steps in the method for anti-occlusion human head three-dimensional positioning based on instance segmentation as described in the first embodiment above are implemented.
[0148] Embodiment 4
[0149] This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps in the method for anti-occlusion human head three-dimensional positioning based on instance segmentation as described in the first embodiment are implemented.
[0150] The steps or modules involved in the above embodiments 2 to 4 correspond to those in embodiment 1. For the specific implementation, please refer to the relevant description of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0151] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for anti-occlusion human head three-dimensional positioning based on instance segmentation, characterized in that: include: Obtain a high-resolution video frame and its corresponding low-resolution video frame; Input the low-resolution video frame into the feature extraction network to obtain shallow features and deep features; Inputting the deep features into a feature fusion detection network to obtain fusion features and low-resolution detection results; the low-resolution detection results include instances with occluded human heads and instances without occluded human heads; Input shallow features and fused features into the feature decoding segmentation network to obtain low-resolution feature maps and low-resolution head segmentation masks; Based on the high-resolution video frame, the unobstructed head instance, the low-resolution feature map and the low-resolution head segmentation mask, the high-resolution unobstructed head instance segmentation result is obtained; Based on the high-resolution video frame, the occluded head instance and the low-resolution head segmentation mask, the segmentation result of the head instance after high-resolution de-occlusion is obtained; specifically: the low-resolution head segmentation mask is up-sampled to obtain the high-resolution head segmentation mask; the occluded head instance is restored to high resolution according to the high-resolution video frame to obtain the bounding box coordinates of the occluded head instance at high resolution; the bounding box coordinates of the occluded head instance at high resolution are expanded, and the high-resolution video frame and the high-resolution head segmentation mask are cropped using the expanded bounding box coordinates to obtain the high-resolution local image and local segmentation mask corresponding to each occluded head instance respectively; the high-resolution local image and local segmentation mask corresponding to each occluded head instance are spliced and sent to the UNet segmentation network to output the occlusion mask in the high-resolution local image corresponding to the occluded head instance; the high-resolution local image corresponding to the occluded head instance is fused with the occlusion mask output in the previous step, and input into the LaMa-based de-occlusion head segmentation network to obtain the segmentation result of the head instance after high-resolution de-occlusion; The high-resolution unobstructed head instance segmentation results and the high-resolution de-occluded head instance segmentation results are input into the three-dimensional positioning module to obtain the three-dimensional coordinates of the head; specifically: the area occupied by each head instance in the image is obtained from the high-resolution unobstructed head instance segmentation results and the high-resolution de-occluded head instance segmentation results, and the distance between the head and the camera in the direction of the camera optical axis is calculated using the square root of the area occupied by the head in the image and the inverse proportional model of the distance between the head and the camera in the direction of the camera optical axis; based on the obtained distance between the head and the camera in the direction of the camera optical axis, the camera intrinsic parameters and the bounding box coordinates of the head, the three-dimensional coordinates of the head in the camera coordinate system are obtained.
2. The method for anti-occlusion human head three-dimensional positioning based on instance segmentation as claimed in claim 1, characterized in that: The low-resolution video frame is input into the feature extraction network to obtain shallow features and deep features, specifically including: the feature extraction network includes 6 network blocks connected in series, each network block outputs a feature map; the three feature maps output by the first three network blocks are shallow features, which are the first, second and third shallow features, respectively, and the three feature maps output by the last three network blocks are deep features; wherein, the first two network blocks each include a Fourier convolution layer, and the last four network blocks each include an image block fusion layer and two visual mamba layers.
3. The method for anti-occlusion human head three-dimensional positioning based on instance segmentation as claimed in claim 1, characterized in that: The feature fusion detection network includes a fast spatial pyramid pooling block, a bidirectional feature fusion block and a detection head; Among them, the fast spatial pyramid pooling block is used to further extract and fuse high-level image features based on the three deep features; the bidirectional feature fusion block is used to fuse features of different scales to obtain the first, second and third fused features corresponding to the deep features; the fused features are respectively sent to their respective detection heads, which include two branches composed of convolutions, one branch is used to output the coordinates of the head bounding box, and the other branch is used to output the head category, where the category includes unoccluded head instances and occluded head instances.
4. The method for anti-occlusion human head three-dimensional positioning based on instance segmentation as claimed in claim 2, characterized in that: The shallow features and fused features are input into the feature decoding segmentation network to obtain a low-resolution feature map and a low-resolution head segmentation mask, specifically including: The feature decoding segmentation network includes three network blocks connected in series and a segmentation head; The first network block takes the first fusion feature as input, passes through an image expansion layer, and then concatenates it with the third shallow feature in the channel dimension and sends it to a visual Mamba layer to output a feature map. ; The second network block starts with As input, after the upsampling layer, it is concatenated with the second shallow layer features and sent to a Fourier convolution layer to output the feature map ; The third network block starts with As input, after the upsampling layer, it is concatenated with the first shallow feature layer and sent to a Fourier convolution layer to output a low-resolution feature map. ; Split the header to As input, after two Fourier convolution layers, the output is a low-resolution head segmentation mask .
5. The method for anti-occlusion human head three-dimensional positioning based on instance segmentation as claimed in claim 1, characterized in that: The method obtains a high-resolution unobstructed head instance segmentation result based on the high-resolution video frame, the unobstructed head instance, the low-resolution feature map and the low-resolution head segmentation mask; specifically includes: Upsampling the low-resolution head segmentation mask and the low-resolution feature map to obtain a high-resolution head segmentation mask and a high-resolution feature map; According to the high-resolution video frame, the unobstructed head instance is restored to high resolution, and the bounding box coordinates of the unobstructed head instance at high resolution are obtained; The high-resolution video frame, the high-resolution head segmentation mask and the high-resolution feature map are cropped using the bounding box coordinates of the unobstructed head instance at high resolution to obtain the high-resolution local image, local feature map and local segmentation mask corresponding to each unobstructed human head instance; The feature map obtained after deformable Fourier convolution of the high-resolution local picture corresponding to each unobstructed head instance is spliced with the local feature map of the corresponding head instance and sent to the deformable Fourier convolution residual block. The output feature is superimposed on the local segmentation mask of the corresponding head instance to obtain the high-resolution unobstructed head instance segmentation result.
6. The method for anti-occlusion human head three-dimensional positioning based on instance segmentation as claimed in claim 1, characterized in that: The high-resolution unobstructed head instance segmentation result and the high-resolution de-occluded head instance segmentation result are input into the 3D positioning module to obtain the 3D coordinates of the head, which also includes: According to the obtained three-dimensional coordinates of the head in the camera coordinate system and the camera extrinsic parameters, the three-dimensional coordinates of the head in the world coordinate system are calculated, and the three-dimensional positioning result of the person in the single camera is obtained.
7. A system for anti-occlusion human head three-dimensional positioning based on instance segmentation, used to execute the method for anti-occlusion human head three-dimensional positioning based on instance segmentation as claimed in claim 1, characterized in that: include: An image acquisition module, used to acquire a high-resolution video frame and its corresponding low-resolution video frame; A low-resolution feature extraction module is used to input low-resolution video frames into a feature extraction network to obtain shallow features and deep features; A low-resolution detection module, used to input deep features into a feature fusion detection network to obtain fusion features and low-resolution detection results; the low-resolution detection results include instances with occluded human heads and instances without occluded human heads; The low-resolution segmentation module is used to input shallow features and fused features into the feature decoding segmentation network to obtain a low-resolution feature map and a low-resolution head segmentation mask; A high-resolution unobstructed segmentation refinement module is used to obtain a high-resolution unobstructed head instance segmentation result based on a high-resolution video frame, an unobstructed head instance, a low-resolution feature map, and a low-resolution head segmentation mask; A high-resolution deocclusion restoration segmentation module is used to obtain a high-resolution deocclusion head instance segmentation result based on a high-resolution video frame, an occluded head instance, and a low-resolution head segmentation mask; The three-dimensional positioning module is used to input the high-resolution unobstructed head instance segmentation result and the high-resolution de-occluded head instance segmentation result into the three-dimensional positioning module to obtain the three-dimensional coordinates of the head.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method for three-dimensional positioning of a human head based on instance segmentation and anti-occlusion are implemented as described in any one of claims 1-6.
9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the method for three-dimensional positioning of a human head based on instance segmentation with anti-occlusion are implemented as described in any one of claims 1-6.
Citation Information
Patent Citations
Image processing method and device, storage medium and equipment
CN110070056A
Trajectory loopback detection optimization method based on generative adversarial network
CN110689562A