Method and system for training lying-posture human body detection model, and lying-posture human body detection method and system

By constructing a lateral human body detection model including FPN network and multi-scale fusion module, the problem of low human body detection accuracy in endoscopic examination scenarios is solved, and high accuracy and applicability of lateral human body detection is achieved.

WO2025130000A1PCT designated stage expired Publication Date: 2025-06-26GUANGZHOU SIDE MEDICAL TECH CO LTD

Patent Information

Application Number
PCT/CN2024/105239
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-07-12
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The existing human detection algorithm cannot be applied to human detection in endoscopic scenarios, resulting in low detection accuracy and inability to provide accurate position transformation data.

Method used

A training method for lying posture human body detection model is provided. By extracting images of lying posture human body from video data of multiple scenes for annotation, a detection model including backbone network, neck network and head network is constructed. The backbone network adopts FPN network and multi-scale fusion module, and the neck network includes context fusion module and attention module.

Benefits of technology

The ability of the lying posture human body detection model to obtain width information is improved. It is suitable for scenes where the width is much higher than the height when lying posture of the human body, suppresses background interference, highlights human body characteristics, significantly improves detection accuracy, and provides accurate position transformation data for endoscopy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024105239_26062025_PF_FP_ABST
    Figure CN2024105239_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of image detection. Specifically disclosed are a method and system for training a lying-posture human body detection model, and a lying-posture human body detection method and system. The method for training a lying-posture human body detection model comprises: extracting from video data an image comprising a lying-posture human body, using the image as a training image, and performing labeling to obtain a real human body frame; and constructing a lying-posture human body detection model, which comprises a backbone network, a neck network and a head network that are connected in sequence, wherein the backbone network comprises a multi-scale fusion module which is embedded in an FPN and at least uses a convolution kernel with the width greater than the height, and the neck network comprises a context fusion module and an attention module. After the lying-posture human body detection model is trained by using the training image, the capability of the lying-posture human body detection model to acquire width information is improved by means of the multi-scale fusion module, the context fusion module and the attention module, such that the interference of a background area can be suppressed to highlight human body features. The present disclosure is applicable to a scenario where there is a lying-posture human body in endoscope examination and a scenario where the background of an examination table is complex in endoscope examination, thereby improving the accuracy of human body detection in endoscope examination.
Need to check novelty before this filing date? Find Prior Art

Description

Training method for prone human body detection model, prone human body detection method and system

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to Chinese patent application number 202311754454.3, filed with the Chinese Patent Office on December 20, 2023, entitled “Training method for prone human body detection model, prone human body detection method and system,” the entire contents of which are incorporated by reference into this application. Technical Field

[0003] The present disclosure belongs to the field of image detection technology, and in particular relates to a training method for a lying human body detection model, a lying human body detection method, and a system. Background Art

[0004] Human endoscopy is a common medical diagnostic method. Capsule gastroscopy, as a non-invasive, painless and non-invasive examination tool, has received widespread attention and application.

[0005] During portable capsule gastroscopy, a camera mounted on the outside of the body is used to monitor the patient's body posture while they are lying down. This facilitates subsequent operations such as determining the correctness of body position changes. Current machine learning-based human detection algorithms are primarily designed for scenes with standing subjects or simple backgrounds. They are not suitable for detecting subjects in a recumbent position during endoscopic examinations, or for complex scenarios where the patient's examination table is covered with bedding or sterile cloths. Consequently, existing human detection algorithms have low accuracy in detecting subjects in endoscopic examinations and are unable to provide accurate body position change data for endoscopic examinations.

[0006] Summary of the Invention

[0007] The purpose of the embodiments of the present disclosure is to provide a training method for a supine human body detection model, a supine human body detection method and a system, aiming to solve the problem that existing human body detection algorithms are not applicable to human body detection in endoscopic examination scenarios, resulting in low accuracy in human body detection in endoscopic examination scenarios.

[0008] To achieve the above objectives, the present disclosure provides the following technical solutions:

[0009] In a first aspect, the present disclosure provides a training method for a lying human body detection model, specifically comprising the following steps:

[0010] Extracting images including lying human bodies from video data of multiple scenes as training images, and labeling the human bodies in the training images to obtain real human body bounding boxes;

[0011] Constructing a recumbent human body detection model, the recumbent human body detection model comprising a backbone network, a neck network, and a head network connected in sequence, the backbone network comprising an FPN network and a multi-scale fusion module embedded in the FPN network and performing a convolution operation using at least a convolution kernel with a width greater than a height, the neck network comprising a context fusion module and an attention module;

[0012] The training image is used to train the lying human body detection model.

[0013] In a second aspect, the present disclosure provides a method for detecting a lying human body, which specifically includes the following steps:

[0014] Collecting images of a human body using an endoscope as images to be detected;

[0015] Inputting the image to be detected into a pre-trained lying human body detection model to obtain human body detection information;

[0016] Wherein, the supine human body detection model is trained by the training method of the supine human body detection model described in the first aspect.

[0017] In a third aspect, the present disclosure provides a training system for a prone human body detection model, specifically comprising:

[0018] a training image acquisition module configured to extract images including lying human bodies from video data of multiple scenes as training images, and to annotate the human bodies in the training images to obtain real human body bounding boxes;

[0019] a recumbent human body detection model construction module, configured to construct a recumbent human body detection model, the recumbent human body detection model comprising a backbone network, a neck network, and a head network connected in sequence, the backbone network comprising an FPN network and a multi-scale fusion module embedded in the FPN network and performing a convolution operation using at least a convolution kernel with a width greater than a height, the neck network comprising a context fusion module and an attention module;

[0020] The lying human body detection model training module is configured to use the training image to train the lying human body detection model.

[0021] In a fourth aspect, the present disclosure provides a lying human body detection system, specifically comprising:

[0022] an image acquisition module to be detected, configured to acquire an image of a human body using an endoscope as an image to be detected;

[0023] a human body detection module configured to input the image to be detected into a pre-trained lying human body detection model to obtain human body detection information;

[0024] Wherein, the supine human body detection model is trained by the training method of the supine human body detection model described in the first aspect.

[0025] In a fifth aspect, the present disclosure provides an electronic device, comprising:

[0026] at least one processor; and

[0027] a memory communicatively connected to the at least one processor; wherein,

[0028] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the training method of the supine human body detection model described in the first aspect of the present disclosure, and / or the supine human body detection method described in the second aspect.

[0029] In a sixth aspect, the present disclosure provides a computer-readable storage medium storing computer instructions, wherein the computer instructions are configured to enable a processor to implement the training method for a supine human body detection model described in the first aspect of the present disclosure, and / or the supine human body detection method described in the second aspect when executed.

[0030] Compared with the prior art, the present invention has the following advantages:

[0031] The supine human body detection model trained by the embodiment of the present disclosure includes a backbone network, a neck network and a head network connected in sequence, wherein the backbone network includes an FPN network and a multi-scale fusion module embedded in the FPN network, which uses at least a convolution kernel with a width greater than a height for convolution operations, and the neck network includes a context fusion module and an attention module, so that the backbone network can perform a convolution operation on the feature map through a convolution kernel with a width greater than a height when extracting features from the training image, thereby obtaining more features of the training image in width, improving the ability of the supine human body detection model to obtain width information, and is suitable for scenarios where the width of the human body is much higher than the height when the human body is supine. In addition, by processing the features extracted by the backbone network through the context fusion module and the attention module, the interference of the background area can be suppressed and the human features can be highlighted. It is suitable for complex scenarios such as the presence of bedding, sterile cloth, etc. on the examination table during endoscopic examination, greatly improving the accuracy of the supine human body detection model in detecting the human body in the endoscopic examination scene, and providing accurate body position transformation data for endoscopic examination.

[0032] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present disclosure.

[0034] FIG1 shows a flow chart of a method for training a lying human body detection model provided by an embodiment of the present disclosure.

[0035] FIG2 shows a schematic diagram of the model structure of a lying human body detection model in one embodiment;

[0036] FIG3 shows a schematic diagram of a multi-scale convolution kernel of a multi-scale fusion module;

[0037] FIG4 shows a schematic diagram of the network structure of the neck network;

[0038] FIG5 shows a flow chart of a lying human body detection method provided by an embodiment of the present disclosure.

[0039] FIG6 shows an application architecture diagram of a training system for a prone human body detection model provided by an embodiment of the present disclosure;

[0040] FIG7 shows an application architecture diagram of a prone human body detection system provided by an embodiment of the present disclosure;

[0041] FIG8 shows an application architecture diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not intended to limit the present disclosure.

[0043] FIG1 shows a flow chart of a method for training a prone human body detection model according to an embodiment of the present disclosure. Specifically, as shown in FIG1 , the method for training a prone human body detection model according to an embodiment of the present disclosure includes the following steps:

[0044] Step S101 : extracting images including lying human bodies from video data of multiple scenes as training images, and annotating the human bodies in the training images to obtain real human body bounding boxes.

[0045] In one embodiment, video data of multiple scenes including a lying human body can be obtained, key frames containing the human body can be extracted from the video data, data expansion processing is performed based on the key frames to obtain expanded images, and the key frames and the expanded images are determined as training images. Each image is further annotated manually or with the assistance of AI (Artificial Intelligence). These annotations are human body bounding boxes, that is, real human body frames containing human body images. The real human body frames have center point coordinates, width and height, and also carry classification labels of human body postures, such as lying down, prone, side-lying, and other classification labels for lying down postures.

[0046] Exemplarily, video data of different human bodies in lying positions in different scenarios can be collected. Exemplarily, video images of human bodies of different heights, weights, ages, and genders on examination tables, beds, or other scenarios can be collected. In the video images, different human bodies change their lying positions or make various movements during the video collection. Key frames containing human bodies can be extracted from the video data. These key frames can effectively represent the posture and movement of the human body in the lying state, such as lying on the left side, half supporting on the left, lying on the right side, half supporting on the right, prone, supine, etc., and the key frames are flipped, rotated, scaled, randomly occluded, randomly blurred, etc. to obtain expanded images. The expanded images and key frames are used as training images, and real human body frames about the boundaries of the human body are added to these training images to obtain more comprehensive and accurate images including human lying positions, providing sufficient and rich samples for the training of the lying human body detection model.

[0047] Step S102: construct a lying human body detection model. The lying human body detection model includes a backbone network, a neck network, and a head network connected in sequence. The backbone network adopts an FPN (Feature Pyramid Networks) network structure, and a multi-scale fusion module (MM) is embedded in the FPN network. The multi-scale fusion module MM uses multiple convolution operations, including a convolution kernel with a width greater than the height for convolution operations. The neck network includes a context fusion module and an attention module.

[0048] As shown in FIG2 , the lying human body detection model of this embodiment includes a backbone network, a neck network, and a head network, wherein the backbone network is used to obtain training images and extract basic features (such as semantic information and position information) from the training images. The output layer of the backbone network is connected to the input layer of the neck network, and the output layer of the neck network is connected to the input layer of the head network. The backbone network is used to extract target feature maps Pi of multiple levels for the training images. The neck network is used to fuse the target feature maps Pi of multiple levels extracted by the backbone network to obtain image features Oi of the training images. The head network is used to output a human body detection frame based on the image features Oi. The output data specifically includes a human body posture classification label (such as cls in FIG2 ), an intersection-and-union ratio between the human body detection frame and a pre-marked real human body frame (such as IOU in FIG2 ), and regression information of the human body detection frame (such as reg in FIG2 ). The regression information reg specifically includes the center coordinates (x p ,y p ), the width w of the human detection frame p , height h p .

[0049] In an optional embodiment, an FPN (Feature Pyramid Network) network structure can be used as the backbone network, and a multi-scale fusion module MM can be embedded in the FPN network structure. The multi-scale fusion module is used to perform multi-scale convolution operations on the multi-level initial feature maps Ci output by the backbone network. Specifically, as shown in Figure 2, the FPN network includes a sub-network that downsamples and extracts features from the bottom layer to the top layer in a forward transmission (extracting initial feature maps Ci of multiple levels), and a sub-network that upsamples and extracts features from the top layer to the bottom layer in a backward transmission embedded in the multi-scale fusion module MM (extracting target feature maps Pi of multiple levels). The multi-scale fusion module MM is used to generate a target feature map Pi of each level, wherein the multi-scale fusion module MM performs a multi-scale convolution operation on the initial feature maps Ci of each level. The multi-scale convolution operation can be a convolution operation of convolution kernels of multiple scales, which can include a standard convolution operation, an asymmetric hole convolution operation, and an asymmetric convolution operation, wherein the asymmetric can be a convolution operation of a convolution kernel with a width greater than the height.

[0050] As shown in Figure 3, a in Figure 3 is a 3×3 standard convolution operation. In the standard convolution operation, the width and height of the convolution kernel are equal, which is used to perform dense sampling of a 3×3 local area of ​​the input feature map. The width and height of the acquired feature map are both 3. b in Figure 3 is a 3×5 asymmetric convolution operation, that is, a convolution with a height h of 3 and a width w of 5. In this embodiment, the width of the convolution kernel of the asymmetric convolution operation is greater than the height. c in Figure 3 is a 3×3 standard hole convolution. In the standard hole convolution, the convolution kernel is provided with a hole and the width is equal to the height. d in Figure 3 is 3×5 asymmetric hole convolution, in which the convolution kernel is provided with holes and the width is greater than the height, wherein the convolution operation without holes can densely sample the local area to obtain high-quality local features, and the hole convolution can realize sparse sampling, can extract farther features, and improve the receptive field. This embodiment extracts features through the asymmetric convolution operation in the multi-scale fusion module, which can improve the ability of the supine human body detection model to extract features of the image in the width direction, thereby being suitable for the scene when the width of the human body is greater than the height when lying down during endoscopic examination.

[0051] As shown in Figure 4, a context fusion module and an attention module can be constructed as a neck network, wherein the context fusion module is used to fuse the target feature maps (Pi, Pi-1, Pi+1) of multiple adjacent levels output by the FPN network to obtain the fusion feature Fi, and the attention module is used to extract the attention feature Ai from the fusion feature Fi, and multiply the fusion feature Fi and the attention feature Ai pixel by pixel as the image feature Oi of the training image.

[0052] As shown in FIG2 , this embodiment is improved based on yolox, and the head network of yolox can be used as the head network of the lying human body detection model. The head network is used to output the predicted human body bounding box when the image feature Oi is input.

[0053] Step S103: using the training image to train the lying human body detection model.

[0054] After constructing the lying human body detection model as shown in Figure 2, during training, the training image can be first input into the FPN network. The FPN network extracts multiple levels of initial feature maps from the training image, and inputs the initial feature maps of each level into the multi-scale fusion module MM. The multi-scale fusion module performs standard convolution, asymmetric convolution, and asymmetric hole convolution operations on the initial feature maps of each level, and connects the features obtained by the convolution operation to obtain the target feature maps of each level, wherein the width of the convolution kernel of the asymmetric convolution and the asymmetric hole convolution is greater than the height.

[0055] As shown in Figure 2, after the training image is input into the FPN network, the semantic features are first extracted by downsampling from the bottom layer to the top layer through the forward transmission sub-network in the FPN network, and the initial feature maps C1, C2, C3, C4 and C5 of multiple levels are obtained. Then, the fine-grained positioning features are extracted by upsampling from the top layer to the bottom layer through the backward transmission sub-network embedded in the multi-scale fusion module MM, and multiple target feature maps P5, P4, and P3 are obtained. Specifically, the multi-scale fusion module MM can be used to perform multi-scale convolution operation on the initial feature map C5 to obtain the target feature map P5, and the target feature map P 5 is upsampled to obtain an upsampled feature map with the same scale as the initial feature map C4, and the upsampled feature map is spliced ​​with the initial feature map C4 and then passed through the multi-scale fusion module MM to obtain the target feature map P4, and then the target feature map P4 is upsampled to obtain an upsampled feature map with the same scale as the initial feature map C3, and the upsampled feature map is spliced ​​with the initial feature map C3 and then passed through the multi-scale fusion module MM to obtain the initial feature map P3. Of course, the above example only takes feature maps of several levels as examples, and in actual applications, feature maps of multiple levels can be set.

[0056] After the backbone network extracts the target feature maps Pi of each layer, the target feature maps Pi of each layer can be input into the context fusion module of the neck network. The context fusion module is used to fuse the target feature maps of the current layer i, the previous layer i-1, and the next layer i+1 to obtain the fused feature Fi of the current layer i. The fused feature Fi is input into the attention module to extract the attention feature Ai and multiplied with the fused feature Fi to obtain the image feature of the training image.

[0057] Specifically, as shown in Figure 4, assuming that the target feature map of the current level i is Pi, the target feature map Pi-1 of the previous level i-1 and the target feature map Pi+1 of the next level i+1 are obtained, wherein the scale of the target feature map Pi is wi×hi×C, wherein wi is the width of the feature map, hi is the height of the feature map, C is the depth, the scale of the target feature map Pi-1 is wi-1×hi-1×C, and the scale of the target feature map Pi+1 is wi+1×hi+1×C. The scales of the target feature map Pi-1 and the target feature map Pi+1 can be adjusted (i.e., the resize operation in Figure 4) to the same scale as the target feature map Pi, that is, after adjusting the target feature map Pi-1 and the target feature map Pi+1 to the same resolution as the target feature map Pi, the target feature map Pi-1 and the target feature map Pi+1 are adjusted. The image Pi and the target feature map Pi+1 are spliced, and the spliced ​​feature map is convolved with a 3×3 convolution layer to obtain the fusion feature Fi. The fusion feature Fi is then processed by a 3×3 convolution layer to obtain the attention feature Ai with a scale of wi×hi×2. After the attention feature Ai is normalized by the softmax function, its first layer is extracted and considered to be the attention feature with a normalized scale of wi×hi×1. The attention feature with a normalized scale of wi×hi×1 is multiplied with the fusion feature Fi to obtain the training image feature Oi, and the image feature Oi is input into the head network to obtain the predicted human body bounding box of the human body in the training image. The loss is calculated based on the predicted human body bounding box and the annotated real human body bounding box, and the model parameters of the lying human body detection model are adjusted according to the loss until the lying human body detection model converges.

[0058] In one embodiment, the loss can be calculated by the following formula: loss = 1 - IOU 2 +λ(w gt -w pred ) 2

[0059] Among them, IOU (Intersection over Union) is the intersection over union ratio of the predicted human body border and the real human body border, λ is the weighting coefficient, and w gt is the width of the real human body frame, w pred is the width of the predicted human body border.

[0060] It can be seen from the above formula that in addition to calculating the intersection-union ratio, the loss in width between the predicted human bounding box and the annotated real human bounding box is also calculated, and a weighted coefficient λ is assigned. When the model parameters are adjusted by this loss, the lying human body detection model pays more attention to the degree of fit between the predicted human bounding box and the real human bounding box in width, that is, the lying human body detection model can better adapt to the changes in the width direction of the human body bounding box in the lying human body scene, that is, it is more suitable for the scene where the human body width is much greater than the height during lying human body detection, thereby improving the ability of the lying human body detection model to detect lying human bodies.

[0061] The supine human body detection model trained in this embodiment includes a backbone network, a neck network and a head network connected in sequence. During training, the backbone network can perform multi-scale convolution fusion operations on the feature map through a convolution kernel with a width greater than the height when extracting features from the training image, thereby obtaining more features of the training image in width, improving the ability of the supine human body detection model to obtain width information, and is suitable for scenarios where the width of the human body is much higher than the height when the human body is supine. In addition, the features extracted by the backbone network are processed by the context fusion module and the attention module, which can suppress the interference of the background area and highlight the human body features. It is suitable for complex scenarios such as the presence of bedding, sterile cloth, etc. on the inspection table during endoscopic examinations, and greatly improves the accuracy of the supine human body detection model in detecting the human body in the endoscopic examination scene, providing accurate body position change data for endoscopic examinations.

[0062] FIG5 shows a flow chart of a lying human body detection method provided by an embodiment of the present disclosure. Specifically, as shown in FIG5 , the lying human body detection method of the embodiment of the present disclosure specifically includes the following steps:

[0063] S201 : collecting images of a human body using an endoscope as images to be detected.

[0064] In one embodiment, a camera can be set outside the human body using an endoscope to capture images of the human body. For example, the patient lies on an examination table in an endoscopy examination room. A camera can be set at a position outside the examination table in the endoscopy examination room and calibrated. Before the endoscope enters the human body, image or video data of the patient on the examination table can be captured. The captured image or video frame sampled from the video data is determined as the image to be detected to detect the body position of the patient on the examination table in the image to be detected.

[0065] S202: Input the image to be detected into a pre-trained lying human body detection model to obtain human body detection information.

[0066] Among them, the supine human body detection model of this embodiment is trained by the training method of the supine human body detection model provided by the embodiment of the present disclosure. After determining the image to be detected, the image to be detected is input into the supine human body detection model to obtain human body detection information, which can include information such as human body posture classification.

[0067] The supine human body detection model used in the supine human body detection method of this embodiment includes a backbone network, a neck network and a head network connected in sequence. After the image to be detected is input into the backbone network, the backbone network can perform a multi-scale convolution fusion operation on the feature map through a convolution kernel with a width greater than the height, thereby obtaining more features in the width of the image to be detected, improving the ability of the supine human body detection model to obtain width information, and is suitable for scenarios where the width of the human body is much higher than the height when the human body is in a supine position. In addition, the features extracted by the backbone network are processed by the context fusion module and the attention module, which can suppress the interference of the background area and highlight the human body features. It is suitable for complex scenarios such as bedding and sterile cloth on the inspection table during endoscopic examination, and greatly improves the accuracy of the supine human body detection model in detecting the human body in the endoscopic examination scene, providing accurate body position change data for endoscopic examination.

[0068] FIG6 shows an application architecture diagram of a training system for a prone human body detection model provided by an embodiment of the present disclosure. The training system for a prone human body detection model of this embodiment includes:

[0069] The training image acquisition module 301 is configured to extract images including lying human bodies from video data of multiple scenes as training images, and annotate the human bodies in the training images to obtain real human body bounding boxes;

[0070] a recumbent human body detection model construction module 302, configured to construct a recumbent human body detection model, the recumbent human body detection model comprising a backbone network, a neck network, and a head network connected in sequence, the backbone network comprising an FPN network and a multi-scale fusion module embedded in the FPN network and performing a convolution operation using at least a convolution kernel with a width greater than a height, the neck network comprising a context fusion module and an attention module;

[0071] The prone human body detection model training module 303 is configured to train the prone human body detection model using training images.

[0072] In an optional embodiment, the training image acquisition module 301 specifically includes:

[0073] a video data acquisition unit configured to acquire video data of a plurality of scenes including a lying human body;

[0074] a key frame extraction unit configured to extract key frames containing a human body from the video data;

[0075] The data expansion unit is configured to perform data expansion processing based on the key frame to obtain an expanded image, and determine the key frame and the expanded image as training images.

[0076] In an optional embodiment, the lying human body detection model building module 302 specifically includes:

[0077] A backbone network construction unit is configured to use an FPN network as a backbone network and embed a multi-scale fusion module in the FPN network, wherein the multi-scale fusion module is configured to perform multiple convolution operations on the feature maps extracted at each level;

[0078] a neck network construction unit configured to construct a context fusion module and an attention module as a neck network, wherein the context fusion module is configured to fuse feature maps of multiple adjacent layers output by the FPN network to obtain a fused feature, and the attention module is configured to extract an attention feature from the fused feature, and multiply the fused feature and the attention feature as an image feature of the training image;

[0079] The head network construction unit is configured to adopt the head network of Yolox as the head network of the lying human body detection model, and the head network is configured to output the predicted human body bounding box when the image features are input.

[0080] In one embodiment, the lying human body detection model training module 303 specifically includes:

[0081] A training image input unit is configured to input a training image into the FPN network, and the FPN network extracts multiple levels of initial feature maps Ci from the training image;

[0082] A convolution operation unit is configured to perform standard convolution, asymmetric convolution, and asymmetric atrous convolution operations on the initial feature map Cn of the last level n through a multi-scale fusion module to obtain a target feature map Pn of the nth level, where n is the number of levels of the FPN network;

[0083] The target feature map extraction unit is configured to upsample the target feature map Pi+1 of the i+1th layer for the i-th layer to obtain an upsampled feature map with the same scale as the initial feature map Ci of the i-th layer, and then perform standard convolution, asymmetric convolution and asymmetric hole convolution operations on the upsampled feature map and the initial feature map Ci of the i-th layer through a multi-scale fusion module to obtain the target feature map Pi of the i-th layer, wherein the width of the convolution kernel of the asymmetric convolution and the asymmetric hole convolution is greater than the height, i≤n-1; the context feature fusion unit is configured to input the target feature map Pi of each layer into the context fusion module of the neck network, and the context fusion module is configured to fuse the target feature maps of the current layer i, the previous layer i-1, and the next layer i+1 to obtain the fused feature Fi of the current layer i;

[0084] An image feature calculation unit is configured to input the fusion feature Fi into the attention module to extract the attention feature Ai and multiply the fusion feature Fi to obtain the image feature of the training image;

[0085] a prediction unit configured to input the image features into the head network to obtain a predicted human bounding box of a human body in the training image;

[0086] A loss calculation unit, configured to calculate a loss based on the predicted human bounding box and the annotated true human bounding box;

[0087] The model parameter adjustment unit is configured to adjust the model parameters of the lying human body detection model according to the loss until the lying human body detection model converges.

[0088] In an optional embodiment, the loss calculation unit specifically includes:

[0089] The loss calculation subunit is configured to calculate the loss using the following formula: loss = 1-IOU 2 +λ(w gt -w pred ) 2

[0090] Among them, IOU is the intersection-over-union ratio of the predicted human body border and the real human body border, λ is the weighting coefficient, and w gt is the width of the labeled real human body border, w pred is the width of the predicted human body border.

[0091] The training system for the supine human body detection model provided in the embodiment of the present disclosure can execute the training method for the supine human body detection model provided in the embodiment of the present disclosure, and has functional modules and beneficial effects corresponding to the execution method.

[0092] FIG7 shows an application architecture diagram of a prone human body detection system provided by an embodiment of the present disclosure. The prone human body detection system of this embodiment includes:

[0093] The image acquisition module 401 is configured to acquire an image of a human body using an endoscope as an image to be detected;

[0094] The human body detection module 402 is configured to input the image to be detected into a pre-trained lying human body detection model to obtain human body detection information;

[0095] The prone human body detection model is trained by the training method for prone human body detection model provided in the embodiment of the present disclosure.

[0096] The prone human body detection system provided by the embodiment of the present disclosure can execute the prone human body detection method provided by the embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0097] FIG8 shows a block diagram of an electronic device 500 that can be used to implement an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0098] As shown in Figure 8, the electronic device 500 includes at least one processor 501 and a memory connected to the at least one processor 501, such as a read-only memory (ROM) 502, a random access memory (RAM) 503, etc., wherein the memory stores a computer program that can be executed by the at least one processor, and the processor 501 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 502 or the computer program loaded from the storage unit 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 can also be stored. The processor 501, ROM 502 and RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0099] Multiple components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0100] The processor 501 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The processor 501 executes the various methods and processes described above, such as the training method of the lying human body detection model and / or the lying human body detection method.

[0101] In some embodiments, the training method of the supine human body detection model, and / or the supine human body detection method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the processor 501, the training method of the supine human body detection model, and / or one or more steps of the supine human body detection method described above can be executed. Alternatively, in other embodiments, the processor 501 can be configured to execute the training method of the supine human body detection model, and / or the supine human body detection method, by any other appropriate means (for example, by means of firmware).

[0102] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0103] Computer programs for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0104] In the context of the present disclosure, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0105] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions of this disclosure can be achieved, and this document is not limited here.

[0106] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure. Industrial Applicability

[0107] Through the above-mentioned training method, supine human body detection method and system, the ability of the supine human body detection model to obtain width information is improved, and it is suitable for scenes where the width of the human body is much larger than the height when the human body is in a supine position. In addition, it can suppress the interference of the background area and highlight the characteristics of the human body. It is suitable for complex scenes such as the presence of bedding, sterile cloth, etc. on the examination table during endoscopic examination. It greatly improves the accuracy of the supine human body detection model in detecting the human body in the endoscopic examination scene, and provides accurate body position change data for endoscopic examination.

Claims

1. A training method for a lying human body detection model, characterized in that: The specific steps include: Extracting images including lying human bodies from video data of multiple scenes as training images, and annotating the human bodies in the training images to obtain real human body bounding boxes; Constructing a lying human body detection model, the lying human body detection model comprising a backbone network, a neck network and a head network connected in sequence, the backbone network comprising an FPN network and a multi-scale fusion module embedded in the FPN network and performing a convolution operation using at least a convolution kernel with a width greater than a height, the neck network comprising a context fusion module and an attention module; The training image is used to train the lying human body detection model.

2. The training method for a prone human body detection model according to claim 1, characterized in that: Extracting images including a lying human body from video data of multiple scenes as training images specifically includes the following steps: Acquire video data of multiple scenes including a lying human body; Extracting key frames containing human bodies from the video data; Data expansion processing is performed based on the key frame to obtain an expanded image, and the key frame and the expanded image are determined as training images.

3. The training method for a lying human body detection model according to claim 1 or 2, characterized in that: The method of constructing a lying human body detection model specifically comprises the following steps: An FPN network is used as a backbone network, and a multi-scale fusion module is embedded in the FPN network, wherein the multi-scale fusion module is used to perform multiple convolution operations on the feature maps extracted at each level; Constructing a context fusion module and an attention module as a neck network, wherein the context fusion module is used to fuse the feature maps of multiple adjacent levels output by the FPN network to obtain a fusion feature, and the attention module is used to extract an attention feature from the fusion feature, and multiply the fusion feature and the attention feature as the image feature of the training image; The head network of Yolox is used as the head network of the lying human body detection model, and the head network is used to output a predicted human body bounding box when image features are input.

4. The training method for a prone human body detection model according to claim 3, characterized in that: The method of using the training image to train the lying human body detection model specifically includes the following steps: The training image is input into the FPN network, and the FPN network extracts multiple levels of initial feature maps Ci from the training image; The initial feature map Cn of the last level n is subjected to standard convolution, asymmetric convolution and asymmetric hole convolution through the multi-scale fusion module to obtain the target feature map Pn of the nth level, where n is the number of levels of the FPN network. For the i-th level, upsample the target feature map Pi+1 of the i+1-th level to obtain an upsampled feature map with the same scale as the initial feature map Ci of the i-th level, concatenate the upsampled feature map with the initial feature map Ci of the i-th level, and perform standard convolution, asymmetric convolution and asymmetric hole convolution operations through a multi-scale fusion module to obtain the target feature map Pi of the i-th level, the width of the convolution kernel of the asymmetric convolution and the asymmetric hole convolution is greater than the height, i≤n-1; Input the target feature map Pi of each level into the context fusion module of the neck network, and the context fusion module is used to fuse the target feature maps of the current level i, the previous level i-1, and the next level i+1 to obtain the fused feature Fi of the current level i; Input the fused feature Fi into the attention module to extract the attention feature Ai and multiply it with the fused feature Fi to obtain the image feature Oi of the training image; Inputting the image feature Oi into the head network to obtain a predicted human body bounding box of the human body in the training image; Calculate the loss based on the predicted human body bounding box and the real human body bounding box; The model parameters of the lying human body detection model are adjusted according to the loss until the lying human body detection model converges.

5. The training method for a lying human body detection model according to claim 3, characterized in that: Calculating the loss based on the predicted human body border and the real human body border specifically includes the following steps: The loss is calculated using the following formula: loss=1-IOU 2 +λ(w gt -w pred ) 2 ; Among them, IOU is the intersection-over-union ratio of the predicted human body border and the real human body border, λ is the weighting coefficient, and w gt is the width of the labeled real human body border, w pred To predict the width of the human body border.

6. A method for detecting a lying human body, characterized in that: include: Collecting images of a human body using an endoscope as images to be detected; Inputting the image to be detected into a pre-trained lying human body detection model to obtain human body detection information; Wherein, the supine human body detection model is trained by the training method for a supine human body detection model described in any one of claims 1-5.

7. A training system for a prone human body detection model, characterized in that: Specifically include: A training image acquisition module is configured to extract images including lying human bodies from video data of multiple scenes as training images, and annotate the human bodies in the training images to obtain real human body frames; a lying human body detection model construction module, configured to construct a lying human body detection model, the lying human body detection model comprising a backbone network, a neck network and a head network connected in sequence, the backbone network comprising an FPN network and a multi-scale fusion module embedded in the FPN network and performing a convolution operation using at least a convolution kernel with a width greater than a height, the neck network comprising a context fusion module and an attention module; The lying human body detection model training module is configured to use the training image to train the lying human body detection model.

8. The training system for the prone human body detection model according to claim 7, characterized in that: The training image acquisition module specifically includes: A video data acquisition unit configured to acquire video data of a plurality of scenes including a lying human body; A key frame extraction unit configured to extract a key frame containing a human body from the video data; The data expansion unit is configured to perform data expansion processing based on the key frame to obtain an expanded image, and determine the key frame and the expanded image as training images.

9. The training system for a prone human body detection model according to claim 7, characterized in that: The lying human body detection model construction module specifically includes: A backbone network construction unit is configured to use an FPN network as a backbone network and embed a multi-scale fusion module in the FPN network, wherein the multi-scale fusion module is configured to perform a plurality of convolution operations on feature maps extracted at each level; A neck network construction unit is configured to construct a context fusion module and an attention module as a neck network, wherein the context fusion module is configured to fuse feature maps of multiple adjacent layers output by the FPN network to obtain a fusion feature, and the attention module is configured to extract an attention feature from the fusion feature, and multiply the fusion feature and the attention feature as an image feature of a training image; The head network construction unit is configured to use the head network of Yolox as the head network of the lying human body detection model, and the head network is configured to output a predicted human body bounding box when inputting image features.

10. The training system for a prone human body detection model according to claim 7, characterized in that: The lying human body detection model training module specifically includes: A training image input unit is configured to input the training image into the FPN network, and the FPN network extracts multiple levels of initial feature maps Ci from the training image; A convolution operation unit is configured to perform standard convolution, asymmetric convolution and asymmetric hole convolution operations on the initial feature map Cn of the last level n through a multi-scale fusion module to obtain a target feature map Pn of the nth level, where n is the number of levels of the FPN network; The target feature map extraction unit is configured to upsample the target feature map Pi+1 of the i+1th layer for the i-th layer to obtain an upsampled feature map of the same scale as the initial feature map Ci of the i-th layer, and then perform standard convolution, asymmetric convolution and asymmetric hole convolution operations through a multi-scale fusion module after splicing the upsampled feature map with the initial feature map Ci of the i-th layer to obtain the target feature map Pi of the i-th layer, wherein the width of the convolution kernel of the asymmetric convolution and the asymmetric hole convolution is greater than the height, i≤n-1; the context feature fusion unit is configured to input the target feature map Pi of each layer into the context fusion module of the neck network, and the context fusion module is configured to fuse the target feature maps of the current layer i, the previous layer i-1, and the next layer i+1 to obtain the fusion feature Fi of the current layer i; An image feature calculation unit is configured to input the fused feature Fi into the attention module to extract the attention feature Ai and multiply it with the fused feature Fi to obtain the image feature of the training image; A prediction unit, configured to input the image features into the head network to obtain a predicted human bounding box of a human body in the training image; A loss calculation unit, configured to calculate a loss based on a predicted human body bounding box and annotated true human body bounding box; The model parameter adjustment unit is configured to adjust the model parameters of the lying human body detection model according to the loss until the lying human body detection model converges.

11. The training system for a prone human body detection model according to claim 10, characterized in that: The loss calculation unit specifically includes: The loss calculation subunit is configured to calculate the loss using the following formula: loss=1-IOU 2 +λ(w gt -w pred ) 2 Among them, IOU is the intersection-over-union ratio of the predicted human body border and the real human body border, λ is the weighting coefficient, and w gt is the width of the labeled real human body border, w pred is the width of the predicted human body border.

12. A lying human body detection system, characterized in that: Specifically include: The image acquisition module to be detected is configured to acquire an image of a human body using an endoscope as an image to be detected; A human body detection module is configured to input the image to be detected into a pre-trained lying human body detection model to obtain human body detection information; Wherein, the supine human body detection model is trained by the training method for a supine human body detection model described in any one of claims 1-5.

13. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the training method for the prone human body detection model described in any one of claims 1-5, and / or the prone human body detection method described in claim 6.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are configured to enable a processor to implement the training method for a prone human body detection model described in any one of claims 1 to 5, and / or the prone human body detection method described in claim 6 when executed.

Citation Information

Patent Citations

  • Target detection method and system based on multi-convolution fusion network

    CN113255589A

  • Target detection method and system, electronic equipment and storage medium

    CN113947154A

  • Target detection method and moving target tracking method using same

    CN114092820A

  • Human body posture estimation model training method, human body posture estimation method, human body posture estimation device and electronic equipment

    CN117037215A

  • Training method of prone position human body detection model, and prone position human body detection method and system

    CN117437697A

Cited By

  • Feature fusion method based on special-shaped convolution kernel

    CN120563989A