Detection system for a disease of a child, detection device and storage medium

By improving the feature fusion module of the YOLOv5 model, images of multiple body parts of children are divided into multiple feature regions and assigned different weight values, which solves the problem of low accuracy caused by small images in children's disease detection and improves the accuracy of detection.

CN115222673BActive Publication Date: 2025-12-16SHENZHEN YUNLING TIANXIA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210751603.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-29
Publication Date
2025-12-16
Estimated Expiration
2042-06-29

AI Technical Summary

Technical Problem

In existing technologies, when detecting diseases in children, the accuracy of existing detection models is low due to the small size of the acquired images, making it impossible to obtain accurate disease detection results.

Method used

An improved detection model based on YOLOv5 is used to detect diseases in multiple images of various body parts of children. The feature fusion module divides the image into multiple feature regions and assigns different weight values ​​to each feature region to improve the accuracy of the detection model.

Benefits of technology

It improves the accuracy of pediatric disease detection, especially in feature extraction from small images and capture of subtle feature information, thus enhancing the accuracy of detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115222673B_ABST
    Figure CN115222673B_ABST
Patent Text Reader

Abstract

The application is suitable for the field of artificial intelligence technology, and provides a detection system, a detection device and a storage medium for a child disease. The system can perform the following operations: multiple images of multiple body parts of a child are collected, and the multiple images at least include an oral cavity image of the child; if the opening and closing degree of the child in the oral cavity image is greater than a preset threshold, a detection model improved based on a YOLOv5 model is used to detect the multiple images, and a detection result is output; wherein the detection model includes an input layer, a bone layer, a neck layer and an output layer, the neck layer includes a feature fusion module, the feature fusion module is used to divide the multiple images into multiple feature regions respectively, and a weight value is assigned to each feature region, and the weight values of each feature region are not completely the same. By using the above system, the accuracy of the detection of the child disease can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The embodiment of the present application belongs to the technical field of artificial intelligence, and particularly relates to a detection system for children's diseases, a detection device and a storage medium. BACKGROUND

[0002] Regular disease detection is helpful to early discovery of various problems in the body and timely medical treatment. In the prior art, images of various parts of the body can be collected, and target detection technology can be used to analyze the images to obtain a disease detection result.

[0003] However, target detection on images can only obtain relatively accurate results when the detection target is relatively large. For some smaller detection targets, the accuracy of the target detection model in the prior art is relatively low, and the output detection result is not accurate enough. For example, when detecting diseases of children, since the image of the part of the body of the child collected as the detection target is relatively small, the existing detection model cannot obtain a relatively accurate disease detection result. SUMMARY

[0004] Therefore, the embodiment of the present application provides a detection system for children's diseases, a detection device and a storage medium, which can improve the accuracy of detection of children's diseases.

[0005] A first aspect of the embodiment of the present application provides a detection system for children's diseases, and the system is applied to perform the following operations:

[0006] Collecting multiple images of multiple parts of the body of a child, wherein the multiple images at least include an oral cavity image of the child;

[0007] If the opening and closing degree of the child in the oral cavity image is greater than a preset threshold, a detection model improved based on a YOLOv5 model is used to perform disease detection on the multiple images, and a detection result is outputted;

[0008] The detection model includes an input layer, a bone layer, a neck layer and an output layer, the neck layer includes a feature fusion module, the feature fusion module is used to divide the multiple images into multiple feature regions respectively, and a weight value is assigned to each feature region, and the weight values of the feature regions are not completely the same.

[0009] A second aspect of the embodiment of the present application provides a detection device for children's diseases, and the device includes a collection module and a detection module, wherein:

[0010] The collection module is used to collect multiple images of multiple parts of the body of a child, wherein the multiple images at least include an oral cavity image of the child;

[0011] The detection module is configured to, if the mouth opening degree of the child in the oral cavity image is greater than a preset threshold, perform disease detection on the multiple images by using a detection model improved based on a YOLOv5 model, and output a detection result.

[0012] The detection model includes an input layer, a bone layer, a neck layer, and an output layer. The neck layer includes a feature fusion module. The feature fusion module is configured to divide the multiple images into multiple feature regions respectively, and assign a weight value to each feature region. The weight values of the feature regions are not completely same.

[0013] A third aspect of the embodiment of the present application provides a detection device including a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor implements the following method when executing the computer program.

[0014] The multiple images of the child at multiple body parts are collected, and the multiple images include at least an oral cavity image of the child.

[0015] If the mouth opening degree of the child in the oral cavity image is greater than a preset threshold, a detection model improved based on a YOLOv5 model is used to perform disease detection on the multiple images, and a detection result is output.

[0016] The detection model includes an input layer, a bone layer, a neck layer, and an output layer. The neck layer includes a feature fusion module. The feature fusion module is configured to divide the multiple images into multiple feature regions respectively, and assign a weight value to each feature region. The weight values of the feature regions are not completely same.

[0017] A fourth aspect of the embodiment of the present application provides a computer readable storage medium storing a computer program. The computer program is executable by a processor to implement the following method.

[0018] The multiple images of the child at multiple body parts are collected, and the multiple images include at least an oral cavity image of the child.

[0019] If the mouth opening degree of the child in the oral cavity image is greater than a preset threshold, a detection model improved based on a YOLOv5 model is used to perform disease detection on the multiple images, and a detection result is output.

[0020] The detection model includes an input layer, a bone layer, a neck layer, and an output layer. The neck layer includes a feature fusion module. The feature fusion module is configured to divide the multiple images into multiple feature regions respectively, and assign a weight value to each feature region. The weight values of the feature regions are not completely same.

[0021] The fifth aspect of the embodiment of the present application provides a computer program product, which, when running on a computer, causes the above computer to execute the following method:

[0022] a plurality of images of a plurality of body parts of a child are collected, and the plurality of images at least include an oral cavity image of the child;

[0023] If the opening and closing degree of the child in the oral cavity image is greater than a preset threshold, a detection model improved based on a YOLOv5 model is used to perform disease detection on the plurality of images, and a detection result is output.

[0024] The detection model includes an input layer, a bone layer, a neck layer, and an output layer, the neck layer includes a feature fusion module, the feature fusion module is used to divide the plurality of images into a plurality of feature regions respectively, and assign a weight value to each feature region, and the weight values of each feature region are not completely the same.

[0025] Compared with the prior art, the embodiment of the present application has the following advantages:

[0026] The embodiment of the present application provides a detection system for child diseases. When the system is used to detect child diseases, a plurality of images of a plurality of body parts of a child can be collected first, and the images at least include an oral cavity image of the child. Before processing the plurality of images, the system can determine whether the opening and closing degree of the child in the collected oral cavity image is greater than a preset threshold. If the opening and closing degree is greater than the preset threshold, the system can use a detection model improved based on a YOLOv5 model to perform disease detection on the plurality of images, and output a detection result. The detection model includes an input layer, a bone layer, a neck layer, and an output layer, and the neck layer includes a feature fusion module. Compared with the traditional YOLOv5 model, the feature fusion module of the detection model in the embodiment of the present application can divide the plurality of images into a plurality of feature regions respectively, and assign different weight values to each feature region according to the importance of each feature region. For example, the weight value corresponding to a feature region with high importance is large, and the weight value corresponding to a feature region with low importance is small. In this way, when the detection model performs disease detection on the images, it can pay more attention to the feature regions with high importance, thereby improving the accuracy of child disease detection. BRIEF DESCRIPTION OF DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creating any creative labor.

[0028] Figure 1 is an operation flow diagram of a disease detection system for detecting a child disease according to an embodiment of the present application;

[0029] Figure 2 is a structure diagram of a detection model improved based on a YOLOv5 model according to an embodiment of the present application;

[0030] Figure 3 is an operation flow diagram of a disease detection system for detecting a child disease according to an embodiment of the present application;

[0031] Figure 4 is an implementation manner of S305 in the operation flow of the disease detection system for detecting a child disease according to an embodiment of the present application;

[0032] Figure 5 is an operation flow diagram of a disease detection system for detecting a child disease according to an embodiment of the present application;

[0033] Figure 6 is an implementation manner of S501 in the operation flow of the disease detection system for detecting a child disease according to an embodiment of the present application;

[0034] Figure 7 is an implementation manner of S502 in the operation flow of the disease detection system for detecting a child disease according to an embodiment of the present application;

[0035] Figure 8 is an example of a disease detection system for detecting a child disease according to an embodiment of the present application;

[0036] Figure 9 is a schematic diagram of a disease detection device according to an embodiment of the present application; and

[0037] Figure 10 is a schematic diagram of a detection device according to an embodiment of the present application. DETAILED DESCRIPTION

[0038] In the following description, specific details are set forth, such as a particular system architecture, techniques, etc., in order to provide a thorough understanding of the present application. However, persons skilled in the art will understand that the present application can be practiced without these specific details. In other instances, well-known structures, devices, circuits, and methods have not been described in detail in order to avoid obscuring the present application.

[0039] It should be understood that the word “comprise” or variations such as “comprises” or “comprising”, when used in this specification and in the accompanying claims, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0040] It should also be understood that the term “and / or” when used in this specification and in the claims which follow, unless otherwise stated, means any conceivable combination of the associated listed items, and includes all possible combinations.

[0041] As used in this specification and in the claims, the term “if’ can be construed to mean “when” or “once” or “in response to determining” or “in response to detecting”, depending on the context. Similarly, the phrase “if it is determined” or “if [a described condition or event] is detected” can be construed to mean “once it is determined” or “in response to determining” or “once [the described condition or event] is detected” or “in response to detecting [the described condition or event]”, depending on the context.

[0042] In addition, in the description of the application and in the claims which follow, the terms “first”, “second”, “third”, etc. are used only to distinguish descriptions, and cannot be understood as indicating or implying relative importance.

[0043] Reference in the specification to “one embodiment” or “some embodiments” means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. The appearances of the phrases “in one embodiment”, “in some embodiments”, “in other embodiments”, “in additional embodiments”, etc. in various places in the specification are not necessarily all referring to the same embodiment, although they can. The terms “comprise”, “comprises”, “comprising”, “include”, “includes”, “including” and the like are synonymous with “containing” or “comprising”, unless otherwise indicated. The terms “comprise”, “comprises”, “comprising”, “include”, “includes”, “including” and the like are synonymous with “containing” or “comprising”, unless otherwise indicated.

[0044] The technical solutions of the application are described below through specific embodiments.

[0045] Referring to Figure 1 , a schematic diagram of an operation process of disease detection using a child disease detection system provided by an embodiment of the application is shown, which can specifically include the following steps:

[0046] S101, a plurality of images of a plurality of body parts of a child are collected, and the plurality of images at least include an oral cavity image of the child.

[0047] The system in this embodiment can be applied in the scenario of detecting diseases of children, such as detecting hand-eye-mouth diseases of children, detecting whether children have hand abrasions, red eye diseases, oral herpes, etc. The system can also detect teeth of children, such as detecting whether children have tooth decay, misaligned teeth, etc.

[0048] The system can be used to detect diseases of children based on the images of each child collected. Therefore, the system can first collect multiple images of different body parts of a child, and multiple images can be collected for each body part. The collected multiple images at least include an oral image of the child, and can also include an eye image, a left hand image, a right hand image, etc. of the child.

[0049] It should be noted that the multiple images can be collected in real time during detection, or can be collected non-real-time. For example, the system can directly take photos of different body parts of a child through an image collection device such as a camera provided by the system to obtain multiple images. Alternatively, the system can also receive multiple images input through other channels, which can be collected and transmitted to the system through electronic devices such as mobile phones, etc.

[0050] The number of images of the body parts of the child collected can be determined according to actual needs. For example, 5 oral images of the child, 6 eye images of the child, 3 left hand images of the child, and 3 right hand images of the child can be collected respectively.

[0051] In S102, if the mouth opening degree of the child in the oral image is greater than a preset threshold, a detection model improved based on a YOLOv5 model is used to detect diseases of the multiple images, and a detection result is output.

[0052] Specifically, before processing the multiple images, the system can calculate the mouth opening degree of the child in the oral image among the multiple images of the multiple body parts of the child collected. The mouth opening degree of the child can be used to represent the degree of opening of the oral cavity of the child. Generally, the mouth opening degree can be quantitatively represented by the ratio between the width and the height of the mouth when the mouth is opened.

[0053] In the embodiments of the present application, if the mouth opening degree of the child in the oral image is greater than a preset threshold, it can be considered that the multiple images of the multiple body parts of the child collected meet the disease detection requirements of the detection model, and the detection model can be used to detect diseases of the multiple images and output a detection result.

[0054] The detection model in the embodiments of the present application can be improved based on a YOLOv5 model. The YOLOv5 model is a single-stage target detection model that can directly predict the category and position of the detection target in one stage. The structure of the YOLOv5 model includes an input layer, a bone layer, a neck layer, and an output layer.

[0055] The YOLOv5 model can be used to detect the oral cavity part of the collected oral cavity image of the child, to obtain a detection frame containing the oral cavity part of the child, and then calculate the mouth opening degree of the child in the detection frame.

[0056] In order to collect clear and wide-range oral cavity images of the child, the mouth opening degree of the child cannot be too small and needs to meet certain value requirements to facilitate detection by the detection model.

[0057] Specifically, a preset threshold of the mouth opening degree of the child that can collect clear and wide-range oral cavity images of the child can be artificially set, such as the preset threshold being 0.8. If the calculated mouth opening degree of the child is greater than the preset threshold, the collected multiple images of the multiple body parts of the child meet the disease detection requirements of the detection model, and the multiple images can be subjected to disease detection. If the calculated mouth opening degree of the child is less than or equal to the preset threshold, the child can be prompted to open the oral cavity, and the operation of collecting the multiple images of the multiple body parts of the child in S101 can be repeated.

[0058] Specifically, in order to collect clear and wide-range images of the body parts of the child, facilitate detection by the detection model, when the collected multiple images of the multiple body parts of the child are not clear or do not include images of the body parts of the child, the child can be prompted to change the posture of collecting images until clear and wide-range images of the body parts of the child are collected.

[0059] The neck layer of the YOLOv5 model includes multiple modules, each module realizing different functions. Among the modules of the neck layer of the YOLOv5 model, there is a Concat module, which can realize the function of feature fusion. When the YOLOv5 model performs target detection, first, the input layer and the bone layer preliminarily process the input detection image, and input the processed image to the neck layer. Then, some modules in the neck layer further process the image and input it to the Concat module. Then, the Concat module performs feature fusion and inputs it to the remaining modules in the neck layer. Next, the remaining modules in the neck layer further process the image and input it to the output layer. Finally, the output layer performs the final processing on the image and outputs the detection result. When the Concat module performs feature fusion, it treats all features equally and assigns the same weight value to all features. However, different features have different importance, and assigning the same weight value to all features results in a lower accuracy of the detection result output by the detection model. Therefore, the Concat module needs to be improved, and different weight values are assigned to different features according to their importance, so that the accuracy of the detection result output by the detection model is improved.

[0060] Specifically, when multiple images are detected for diseases, a detection model improved based on a YOLOv5 model can be used to detect diseases in the multiple images and output detection results. The structure of the detection model improved based on the YOLOv5 model is shown in Figure 2 The structure of the detection model includes an input layer, a bone layer, a neck layer, and an output layer. Moreover, the detection model improved based on the YOLOv5 model replaces the Concat module of the neck layer of the YOLOv5 model with a feature fusion module. The feature fusion module divides the collected multiple images into multiple feature regions, can extract the features of small-range images, and obtain fine feature information, thereby improving the accuracy of child disease detection. The feature fusion module assigns a weight value to each feature region, and according to the importance of each feature region, assigns different weight values to each feature region. The weight value corresponding to the feature region with high importance is large, and the weight value corresponding to the feature region with low importance is small. When the detection model detects diseases in images, it can pay more attention to the feature region with high importance, thereby improving the accuracy of child disease detection.

[0061] In this embodiment, before the detection model detects diseases, the opening and closing degree of the child's mouth in the collected oral image is calculated, and a clear and wide-range oral image of the child can be collected to facilitate detection by the detection model. The detection model improved based on the YOLOv5 model is used to detect diseases in multiple images, and the collected multiple images are divided into multiple feature regions, the features of small-range images can be extracted, fine feature information can be obtained, and the accuracy of child disease detection can be improved. Moreover, according to the importance of each feature region, different weight values are assigned to each feature, and more attention can be paid to the feature region with high importance, thereby improving the accuracy of child disease detection.

[0062] Referring to Figure 3 , a schematic diagram of an operation process for calculating the opening and closing degree of the mouth by a child disease detection system provided in an embodiment of the present application is shown, which can specifically include the following steps:

[0063] S301, multiple images of multiple body parts of a child are collected, and the multiple images at least include oral images of the child.

[0064] Since S301 in this embodiment is similar to S101 in the foregoing embodiments, they can participate in each other, and this embodiment will not be repeated here.

[0065] S302, multiple detection boxes in each oral image are determined, and the multiple detection boxes are all image regions containing the oral part of the child.

[0066] Specifically, the YOLOv5 model can be used to detect the collected oral cavity image of the child to obtain a plurality of detection boxes containing the oral cavity part of the child. Each oral cavity image contains a plurality of detection boxes, and each detection box is an image region containing the oral cavity part of the child.

[0067] S303, respectively calculating the mouth opening degree of the child in each detection box.

[0068] Specifically, the coordinates of the four corners of each detection box can be obtained, the width and height of each detection box are determined according to the coordinates of the four corners of each detection box, and the mouth opening degree of the child in each detection box is determined according to the ratio of the width and height of each detection box.

[0069] S304, determining the mouth opening degree of the child in the oral cavity image according to the mouth opening degree of the child in each detection box.

[0070] When the YOLOv5 model is used for target detection, a non-maximum suppression algorithm is often used to determine one of the plurality of detection boxes as a detection result according to a threshold value determined by a person. However, when one of the plurality of detection boxes is determined as a detection result, there is obvious redundancy, which leads to the final determined detection box containing not only the image region of the detection target but also the image region of the non-detection target, resulting in a large error and low accuracy of the final detection result. It can be concluded that the detection box determined according to the threshold value determined by a person has a large error and low accuracy in calculating the mouth opening degree of the child in the detection box. Therefore, it is necessary to improve the determination of the mouth opening degree of the child in the oral cavity image according to the non-maximum suppression algorithm to improve the accuracy of the mouth opening degree of the child in the oral cavity image.

[0071] Specifically, when determining the mouth opening degree of the child in the oral cavity image, the mouth opening degree of the child in each detection box can be calculated first, and then the mouth opening degree of the child in the oral cavity image can be determined according to the average of the mouth opening degrees of the child in the plurality of detection boxes. Compared with the threshold value determined by the detection box, the average of the mouth opening degrees of the child in the plurality of detection boxes has a smaller error and higher accuracy.

[0072] Specifically, the child disease detection system can also determine the mouth opening degree of the child in the oral cavity image by using the following formula:

[0073]

[0074] wherein R is the mouth opening degree of the child, n is the number of detection boxes, W m is the width of the mth detection box, and H mK1 is a correction value of the width of the detection frame, K2 is a correction value of the height of the detection frame, and β is a positive term for correcting the opening and closing degree of the child's mouth.

[0075] In the above formula, K1 can correct the redundancy of the detection frame in the longitudinal direction, K2 can correct the redundancy of the detection frame in the transverse direction, and β is the ratio of the width to the height of the oral cavity when the oral cavity is not open. When calculating the opening and closing degree of the child's mouth, the ratio of the width to the height of the oral cavity when the oral cavity is not open is subtracted, which can correct the opening and closing degree of the child's mouth as a whole. Therefore, the above formula can correct the redundancy of the detection frame, focus on the influence of the width and height of each detection frame on the opening and closing degree of the child's mouth in the finally determined oral cavity image, and improve the accuracy of the opening and closing degree of the child's mouth.

[0076] In S305, if the opening and closing degree of the child's mouth in the oral cavity image is greater than a preset threshold, a detection model improved based on a YOLOv5 model is used to detect diseases in the multiple images, and a detection result is output.

[0077] The structure of the detection model improved based on the YOLOv5 model is as shown in Figure 2 The neck layer of the detection model improved based on the YOLOv5 model can further include a pooling module.

[0078] In a possible implementation manner of the embodiments of the present application, as shown in Figure 4 The step of using the detection model improved based on the YOLOv5 model to detect diseases in the multiple images in S305 and outputting the detection result can specifically include the following steps:

[0079] S3051, the feature map output by the previous module of the pooling module is split into a plurality of first feature maps.

[0080] S3052, the feature map output by the previous module of the pooling module is pooled to obtain a second feature map.

[0081] S3053, the plurality of first feature maps and the second feature map are superimposed to obtain an output feature map of the neck layer, and the output feature map is used for feature calculation by an output layer to obtain a detection result.

[0082] When the detection system detects diseases in the multiple images, due to a large amount of calculation, in the output feature map of the neck layer, smaller features are not very obvious and are easily ignored. Therefore, the output feature map of the neck layer needs to be improved.

[0083] Specifically, the feature map output by the previous module of the pooling module can be split into a plurality of first feature maps, the feature map output by the previous module of the pooling module is pooled to obtain a second feature map, the plurality of first feature maps obtained by splitting and the second feature map obtained by pooling are superimposed to obtain an output feature map of a neck layer, and the output feature map of the neck layer is transmitted to a transmission layer of the detection model, and the transmission layer further performs feature calculation to obtain a detection result.

[0084] In the embodiment, when determining the mouth opening degree of the child in the oral cavity image, the mouth opening degree of the child in each detection frame is calculated respectively, and the mean value of the mouth opening degrees of the child in the plurality of detection frames is used to determine the mouth opening degree of the child in the oral cavity image. The influence of the width and height of each detection frame on the finally determined mouth opening degree of the child in the oral cavity image can be considered, and the finally obtained mouth opening degree of the child has smaller error and higher accuracy, which is convenient for collecting clear and wide-range oral cavity images of children and convenient for the detection model to detect. By superimposing the plurality of first feature maps obtained by splitting and the second feature map obtained by pooling, the output feature map of the neck layer is obtained. Compared with directly pooling to obtain the output feature map of the neck layer, smaller features can be noticed, and the output feature map of the neck layer can retain some more detailed information.

[0085] Referring to Figure 5 , an operation flow diagram for determining the weight value of a feature region by using the detection system for child diseases provided in the embodiments of the present application is shown, and the operation flow diagram can specifically include the following steps:

[0086] In S501, for any feature region, a cyclic focusing mechanism is used to assign a first weight value to the feature region.

[0087] When the detection model observes an image to extract observation information, the focusing mechanism can observe only a certain feature region of the image each time, and can change the observed feature region after observing the feature region. The cyclic focusing mechanism can cyclically observe a certain feature region, so that all feature regions of the image are observed.

[0088] Specifically, the feature fusion module of the improved detection model based on the YOLOv5 model can divide a plurality of images output by a previous module of the feature fusion module into a plurality of feature regions, and each of the plurality of images is divided into a plurality of feature regions. The feature fusion module can include a cyclic focusing mechanism and an effective channel attention mechanism, and each feature region is assigned a weight value.

[0089] Specifically, for any feature region, the feature fusion module can use a cyclic focusing mechanism to assign a first weight value to the feature region. First, observe a certain feature region to extract observation information of the feature region, and assign a first weight value to the feature region according to the observation information of the feature region. Next, replace the observed feature region, and repeat the operation of assigning a first weight value to the observed feature region until all feature regions are observed and assigned a first weight value. Therefore, for any feature region, the feature fusion module uses a cyclic focusing mechanism to assign a first weight value to the feature region, and by observing a certain feature region to extract observation information of the feature region, the features of a small range of images can be extracted to obtain fine feature information, thereby improving the accuracy of child disease detection.

[0090] Specifically, when observing multiple feature regions, the order in which the multiple feature regions are observed can be from the first feature region in the top left corner of the image, then from left to right, until all feature regions in the first row of the image from top to bottom are observed, and then from left to right, the feature regions in the second row of the image from top to bottom are observed, until all rows of feature regions are observed.

[0091] S502, using an effective channel attention mechanism to assign a second weight value to the feature region.

[0092] Specifically, for any feature region, the effective channel attention mechanism of the feature fusion module can assign a second weight value to the feature region. Moreover, each feature region is assigned a weight value that is not completely the same, the weight value corresponding to a feature region with high importance is large, and the weight value corresponding to a feature region with low importance is small, which can pay more attention to the feature region with high importance, thereby improving the accuracy of child disease detection.

[0093] S503, determining the weight value of the feature region according to the first weight value and the second weight value.

[0094] Specifically, when determining the weight value of a certain feature region, the matrix of the feature region can be multiplied by the first weight value of the feature region, and then multiplied by the second weight value of the feature region.

[0095] In this embodiment, for any feature region, the feature fusion module uses a cyclic focusing mechanism to assign a first weight value to the feature region, and by observing a certain feature region to extract observation information of the feature region, the features of a small range of images can be extracted to obtain fine feature information, thereby improving the accuracy of child disease detection. For any feature region, the effective channel attention mechanism of the feature fusion module can assign a second weight value to the feature region, and each feature region is assigned a weight value that is not completely the same, which can pay more attention to the feature region with high importance, thereby improving the accuracy of child disease detection.

[0096] In a possible implementation of the embodiment of the present application, as shown in Figure 6 The step S501 of assigning the first weight value to the feature region by using the loop focusing mechanism for any feature region can include the following steps S5011-S5016.

[0097] S5011, determining a current observation position of a current feature region.

[0098] Specifically, the feature fusion module divides the multiple images output by the previous module of the module into multiple feature regions, determines one of the multiple feature regions as the current feature region, and can determine the current observation position of the current feature region according to the current feature region. The current feature region can be represented by a matrix, and the current observation position can be represented by coordinates. The current feature region is in one-to-one correspondence with the current observation position of the current feature region.

[0099] S5012, based on the current observation position, cyclically extracting current observation information of the current feature region, the current observation information including position information and texture information of the current feature region.

[0100] Specifically, the current observation information of the current feature region can be cyclically extracted based on the current observation position of the current feature region, and the extracted current observation information includes the position information and the texture information of the current feature region. The size of the extracted current observation information can represent the importance of the position information and the texture information of the current feature region.

[0101] S5013, determining a first state value of a current loop according to the current observation information and a first state value of a previous loop, the first state value of the previous loop being determined by the observation information extracted in the previous loop.

[0102] Specifically, the first state value of the current loop is determined by the current observation information and the first state value of the previous loop. Since the size of the current observation information can represent the importance of the position information and the texture information of the current feature region, the size of the first state value of the current loop can represent the importance of the position information and the texture information of the current feature region.

[0103] S5014, determining a first weight value according to the first state value of the current loop.

[0104] Specifically, the first weight value of the current feature region is determined by the first state value of the current loop, and the first weight value of the current feature region is assigned to the current feature region.

[0105] The Concat module of the YOLOv5 model does not divide the multiple images output by the previous module of the module into multiple feature regions, and cannot extract the features of small-range images. The Concat module assigns the same weight to all features, but different features have different contributions to the output results of the detection model. Assigning the same weight to all features can make the accuracy of the detection model not high enough.

[0106] In the process of disease detection by the detection model improved based on the YOLOv5 model, the size of the first state value of the current cycle can represent the importance of the position information and the texture information of the current feature region, and then the first weight value of the current feature region can represent the importance of the position information and the texture information of the current feature region. The higher the importance of the position information and the texture information of the current feature region, the greater the weight value corresponding to the current feature region, and the lower the importance of the position information and the texture information of the current feature region, the smaller the weight value corresponding to the current feature region, which can improve the accuracy of the detection model. In addition, by observing the current feature region to extract observation information of the feature region to determine the first weight value of the current feature region, the features of small-range images can be extracted, and fine feature information can be obtained, which is assigned to the current feature region in the form of the first weight value, thereby improving the accuracy of child disease detection.

[0107] S5015, determining the second state value of the current cycle according to the first state value of the current cycle and the second state value of the previous cycle.

[0108] Specifically, the second state value of the current cycle is determined by the first state value of the current cycle and the second state value of the previous cycle.

[0109] S5016, determining the next observation position according to the second state value of the current cycle, the next observation position being an observation position of a next feature region of the current feature region.

[0110] Specifically, the next observation position is determined by the second state value of the current cycle. In the next observation position, the next feature region is observed, and steps S5011-S5016 are repeated until all feature regions are assigned with the first weight value.

[0111] In a possible implementation manner of the embodiment of the present application, as shown in Figure 7 The second weight value is assigned to the feature region by adopting the effective channel attention mechanism in S502, which can specifically include the following steps S5021-S5024:

[0112] S5021, performing global average pooling on the feature region to obtain a feature vector for representing the feature region.

[0113] Specifically, global average pooling can be performed on the feature region to obtain a global receptive field, and a feature vector for representing the feature region is obtained.

[0114] S5022, a fast one-dimensional convolution calculation is performed on the feature vector to capture local cross-channel interaction information, and the convolution size of the fast one-dimensional convolution calculation is proportional to the channel dimension of the global average pooling, and the coverage of the local cross-channel interaction information is determined by the convolution size of the fast one-dimensional convolution calculation.

[0115] Specifically, a fast one-dimensional convolution calculation can be performed on the feature vector obtained by the global average pooling to capture local cross-channel interaction information. The coverage of the local cross-channel interaction information is determined by the convolution size of the fast one-dimensional convolution calculation. The convolution size of the fast one-dimensional convolution is related to the channel dimension of the global average pooling, and is proportional to the channel dimension of the global average pooling. The fast one-dimensional convolution can realize local cross-channel interaction without dimension reduction, and can avoid the influence brought by channel dimension reduction. In the local area, there is appropriate cross-channel interaction, which can improve the performance of the detection model.

[0116] S5023, determining channel attention according to the local cross-channel interaction information.

[0117] Specifically, the channel attention is determined according to the local cross-channel interaction information captured by the fast one-dimensional convolution calculation. The channel attention can be a diagonal matrix, and each channel corresponds to a channel attention. According to the local cross-channel interaction without dimension reduction, the influence brought by channel dimension reduction can be avoided, so that effective channel attention can be learned, and the performance of the detection model can be improved.

[0118] S5024, giving a second weight value to the feature region based on the channel attention.

[0119] Specifically, the second weight value is given to the feature region based on the channel attention determined according to the local cross-channel interaction information.

[0120] In the embodiment, the first weight value of the current feature region determined based on the position information and the texture information of the current feature region represents the importance of the position information and the texture information of the current feature region, which can improve the accuracy of the detection model. By observing the feature region, the features of a small range of images are extracted to obtain fine feature information, which is given to the current feature region in the form of the first weight value, which can improve the accuracy of the detection of children's diseases. Based on the fast one-dimensional convolution, the local cross-channel interaction information without dimension reduction is obtained, which can avoid the influence brought by channel dimension reduction, so that effective channel attention can be learned, and the performance of the detection model can be improved.

[0121] In order to facilitate understanding, the process of detecting children's diseases by applying the children's disease detection system will be introduced below with a specific example.

[0122] Referring to Figure 8 In the application, when the child disease detection system is used to detect the child disease, the image of the body part of the child is first collected, and it is judged whether the child stretches out the hand. If the child does not stretch out the hand, the detection system prompts the child to stretch out the hand, and the operation of judging whether the child stretches out the hand is repeated. If the child stretches out the hand, it is judged whether the child opens the mouth. Whether the child opens the mouth can be judged by calculating the opening degree of the mouth of the child. The specific process of calculating the opening degree of the mouth of the child can be referred to the introduction of the foregoing system embodiment part, and will not be repeated here.

[0123] If the opening degree of the mouth of the child is greater than the threshold value, it is determined that the child opens the mouth, and if the opening degree of the mouth of the child is less than or equal to the threshold value, it is determined that the child does not open the mouth. If the child does not open the mouth, the detection system prompts the child to open the mouth, and the operation of judging whether the child opens the mouth is repeated. If the child opens the mouth, the detection model improved based on the YOLOv5 model is used for disease detection to judge whether the child is ill. The specific process of using the detection model improved based on the YOLOv5 model for disease detection can be referred to the introduction of the foregoing system embodiment part, and will not be repeated here.

[0124] If the child is not ill, the detection system displays normal and ends the detection. If the child is ill, the detection system displays the detected child disease and ends the detection.

[0125] Referring to Figure 9 , a schematic diagram of a child disease detection device provided by an embodiment of the application is shown, the device 900 can include an acquisition module 901 and a detection module 902, wherein:

[0126] The acquisition module 901 is configured to acquire a plurality of images of a plurality of body parts of a child, and the plurality of images at least include an oral cavity image of the child.

[0127] The detection module 902 is configured to, if the opening degree of the mouth of the child in the oral cavity image is greater than a preset threshold value, use a detection model improved based on a YOLOv5 model to perform disease detection on the plurality of images, and output a detection result.

[0128] The detection model includes an input layer, a bone layer, a neck layer and an output layer, the neck layer includes a feature fusion module, the feature fusion module is configured to divide the plurality of images into a plurality of feature regions respectively, and assign a weight value to each feature region, and the weight values of each feature region are not completely same.

[0129] The child disease detection device can further include:

[0130] A detection frame determination module is configured to determine a plurality of detection frames in each oral cavity image, and the plurality of detection frames are all image regions containing the oral cavity part of the child.

[0131] a calculation module configured to calculate the opening-closing degree of the child in each detection frame respectively;

[0132] a determination opening-closing degree module configured to determine the opening-closing degree of the child in the oral cavity image according to the opening-closing degree of the child in each detection frame.

[0133] The device can determine the opening-closing degree of the child in the oral cavity image according to the following formula:

[0134]

[0135] wherein, R is the opening-closing degree of the child, n is the number of detection frames, W m is the width of the mth detection frame, H m is the height of the mth detection frame, K1 is the correction value of the width of the detection frame, K2 is the correction value of the height of the detection frame, and β is a regularization term for correcting the opening-closing degree of the child.

[0136] The feature fusion module in the detection model can include a recurrent focus mechanism and an effective channel attention mechanism, and the detection module 902 can be further configured to call the feature fusion module to assign a weight value to each feature region. When assigning the weight value to each feature region, the feature fusion module is specifically configured to: for any feature region, assign a first weight value to the feature region by using the recurrent focus mechanism; assign a second weight value to the feature region by using the effective channel attention mechanism; and determine the weight value of the feature region according to the first weight value and the second weight value.

[0137] In the embodiments of the present application, the feature fusion module can be further configured to:

[0138] determine a current observation position of the current feature region; cyclically extract current observation information of the current feature region based on the current observation position, wherein the current observation information includes position information and texture information of the current feature region; determine a first state value of the current cycle according to the current observation information and a first state value of a previous cycle, wherein the first state value of the previous cycle is determined based on observation information extracted in the previous cycle; and determine the first weight value according to the first state value of the current cycle.

[0139] In the embodiments of the present application, the feature fusion module can be further configured to: determine a second state value of the current cycle according to the first state value of the current cycle and a second state value of the previous cycle; and determine a next observation position according to the second state value of the current cycle, wherein the next observation position is an observation position of a next feature region of the current feature region.

[0140] The feature fusion module can also be configured to: perform global average pooling on the feature region to obtain a feature vector for representing the feature region; and perform fast one-dimensional convolution calculation on the feature vector to capture local cross-channel interaction information, wherein a convolution size of the fast one-dimensional convolution calculation is proportional to a channel dimension of the global average pooling, and a coverage range of the local cross-channel interaction information is determined by the convolution size of the fast one-dimensional convolution calculation; and determine channel attention according to the local cross-channel interaction information, and assign a second weight value to the feature region based on the channel attention.

[0141] In the embodiments of the present application, the neck layer of the detection model further comprises a pooling module; the pooling module is the last module of the neck layer; and the pooling module is specifically configured to:

[0142] split a feature map output by a previous module of the pooling module into a plurality of first feature maps; perform pooling on the feature map output by the previous module of the pooling module to obtain a second feature map; and stack the plurality of first feature maps and the second feature map to obtain an output feature map of the neck layer, wherein the output feature map is used for feature calculation by an output layer to obtain a detection result.

[0143] Referring to Figure 10 , a schematic diagram of a detection device provided in an embodiment of the present application is shown. As shown in Figure 10 , the detection device 1000 in the embodiment of the present application comprises a processor 1010, a memory 1020, and a computer program 1021 stored in the memory 1020 and executable on the processor 1010. The processor 1010 implements the steps in the various embodiments of the detection system for child diseases when executing the computer program 1021, such as the steps S101-S102 shown in Figure 1 .

[0144] For example, the computer program 1021 can be divided into one or more modules / units, which are stored in the memory 1020 and executed by the processor 1010 to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of completing a specific function, which can be used to describe the execution process of the computer program 1021 in the detection device 1000. For example, the computer program 1021 can be divided into an acquisition module and a detection module, and the specific functions of the modules are as follows:

[0145] The acquisition module is configured to acquire a plurality of images of a plurality of body parts of a child, and the plurality of images at least include an oral cavity image of the child.

[0146] The detection module is configured to, if the opening and closing degree of the child's mouth in the oral cavity image is greater than a preset threshold, perform disease detection on the plurality of images by using a detection model improved based on a YOLOv5 model, and output a detection result.

[0147] The detection model comprises an input layer, a bone layer, a neck layer and an output layer, the neck layer comprises a feature fusion module, the feature fusion module is configured to divide the plurality of images into a plurality of feature regions respectively, and assign a weight value to each feature region, and the weight values of the feature regions are not completely same.

[0148] The detection device 1000 can be an electronic device for implementing the detection system of the child disease in the foregoing embodiments. The detection device 1000 can be a desktop computer, a cloud server or the like. The detection device 1000 can include, but is not limited to, a processor 1010 and a memory 1020. Those skilled in the art can understand that the detection device 1000 can include more or fewer components than those shown, or combine some components, or include different components, for example, the detection device 1000 can also include an input / output device, a network access device, a bus and the like. Figure 10 The detection device 1000 is only an example and does not constitute a limitation on the detection device 1000, and can include more or fewer components than those shown, or combine some components, or include different components, for example, the detection device 1000 can also include an input / output device, a network access device, a bus and the like.

[0149] The processor 1010 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0150] The memory 1020 can be an internal storage unit of the detection device 1000, for example, a hard disk or a memory of the detection device 1000. The memory 1020 can also be an external storage device of the detection device 1000, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card and the like equipped on the detection device 1000. Further, the memory 1020 can include both the internal storage unit and the external storage device of the detection device 1000. The memory 1020 is used to store the computer program 1021 and other programs and data required by the detection device 1000. The memory 1020 can also be used to temporarily store data that has been output or will be output.

[0151] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the following method.

[0152] a plurality of images of a plurality of body parts of a child are collected, and the plurality of images at least include an oral cavity image of the child;

[0153] If the mouth opening degree of the child in the oral cavity image is greater than a preset threshold, a detection model improved based on a YOLOv5 model is used to perform disease detection on the plurality of images, and a detection result is output.

[0154] The detection model includes an input layer, a bone layer, a neck layer and an output layer, the neck layer includes a feature fusion module, the feature fusion module is used to divide the plurality of images into a plurality of feature regions respectively, and a weight value is assigned to each feature region, and the weight values of the feature regions are not completely same.

[0155] The embodiment of the present application also provides a computer program product, when the computer program product runs on a computer, so that the computer executes the following method:

[0156] a plurality of images of a plurality of body parts of a child are collected, and the plurality of images at least include an oral cavity image of the child;

[0157] If the mouth opening degree of the child in the oral cavity image is greater than a preset threshold, a detection model improved based on a YOLOv5 model is used to perform disease detection on the plurality of images, and a detection result is output.

[0158] The detection model includes an input layer, a bone layer, a neck layer and an output layer, the neck layer includes a feature fusion module, the feature fusion module is used to divide the plurality of images into a plurality of feature regions respectively, and a weight value is assigned to each feature region, and the weight values of the feature regions are not completely same.

[0159] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A system for detecting a childhood disease, characterized by, The system comprises a detection device, which performs the following operations: Collecting multiple images of multiple body parts of a child, wherein the multiple images at least include an oral cavity image of the child; If the mouth opening degree of the child in the oral cavity image is greater than a preset threshold, a detection model improved based on a YOLOv5 model is used to perform disease detection on the multiple images, and a detection result is output; The detection model comprises an input layer, a bone layer, a neck layer and an output layer, the neck layer comprises a feature fusion module, the feature fusion module is used to divide the multiple images into multiple feature regions respectively, and assign a weight value to each feature region, and the weight values of each feature region are not completely the same; The feature fusion module comprises a cyclic focusing mechanism and an effective channel attention mechanism, and the feature fusion module is used to assign a weight value to each feature region, specifically comprising: For any feature region, the cyclic focusing mechanism is used to assign a first weight value to the feature region; The effective channel attention mechanism is used to assign a second weight value to the feature region; According to the first weight value and the second weight value, the weight value of the feature region is determined; The cyclic focusing mechanism is used to assign a first weight value to the feature region, comprising: Determine the current observation position of the current feature region; Based on the current observation position, the current observation information of the current feature region is cyclically extracted, and the current observation information includes the position information and the texture information of the current feature region; According to the current observation information and the first state value of the previous cycle, the first state value of the current cycle is determined, and the first state value of the previous cycle is determined by the observation information extracted by the previous cycle; According to the first state value of the current cycle, the first weight value is determined; The effective channel attention mechanism is used to assign a second weight value to the feature region, comprising: Global average pooling is performed on the feature region to obtain a feature vector for representing the feature region; Fast one-dimensional convolution calculation is performed on the feature vector to capture local cross-channel interaction information, the convolution size of the fast one-dimensional convolution calculation is proportional to the channel dimension of the global average pooling, and the coverage range of the local cross-channel interaction information is determined by the convolution size of the fast one-dimensional convolution calculation; According to the local cross-channel interaction information, a channel attention is determined; Based on the channel attention, the second weight value is assigned to the feature region.

2. The system of claim 1, wherein, The application of the system further performs the following operations: Determine multiple detection boxes in each of the oral cavity images, wherein each of the multiple detection boxes is an image region containing the oral cavity part of the child; Calculate the mouth opening degree of the child in each of the detection boxes respectively; According to the mouth opening degree of the child in each of the detection boxes, the mouth opening degree of the child in the oral cavity image is determined.

3. The system of claim 2, wherein, The system determines the mouth opening degree of the child in the oral cavity image by using the following formula: wherein, is an opening-closing degree of the child, is a number of the detection frame, is a width of the detection frame, is a height of the detection frame, is a correction value of the width of the detection frame, is a correction value of the height of the detection frame, is a correction value of the width of the detection frame, is a correction value of the height of the detection frame, is a regularization term for correcting the opening-closing degree of the child.

4. The system of claim 1, wherein, After determining the first weight value according to the first state value of the current cycle, further comprising: determining a second state value of the current cycle according to the first state value of the current cycle and a second state value of the previous cycle; determining a next observation position according to the second state value of the current cycle, the next observation position being an observation position of a next feature region of the current feature region.

5. The system of any one of claims 1-3, wherein, The neck layer further comprises a pooling module; the pooling module is the last module of the neck layer; the system is further applied to perform the following operations: splitting a feature map output by a previous module of the pooling module into a plurality of first feature maps; pooling the feature map output by the previous module of the pooling module to obtain a second feature map; stacking the plurality of first feature maps and the second feature map to obtain an output feature map of the neck layer, the output feature map being used for feature calculation by the output layer to obtain the detection result.

6. A detection device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor implements the following method when executing the computer program: collecting multiple images of multiple body parts of a child, the multiple images comprising at least an oral cavity image of the child; if the opening and closing degree of the child's mouth in the oral cavity image is greater than a preset threshold, using a detection model improved based on a YOLOv5 model to perform disease detection on the multiple images and output a detection result; The detection model comprises an input layer, a bone layer, a neck layer and an output layer, the neck layer comprises a feature fusion module, and the feature fusion module is configured to divide the multiple images into multiple feature regions respectively and assign a weight value to each feature region, and the weight values of the feature regions are not completely the same; The feature fusion module comprises a cycle focus mechanism and an effective channel attention mechanism, and the feature fusion module is configured to assign a weight value to each feature region, specifically including: for any feature region, the cycle focus mechanism is used to assign a first weight value to the feature region; the effective channel attention mechanism is used to assign a second weight value to the feature region; the weight value of the feature region is determined according to the first weight value and the second weight value; The cycle focus mechanism is used to assign a first weight value to the feature region, including: determining a current observation position of a current feature region; based on the current observation position, cyclically extracting current observation information of the current feature region, the current observation information comprising position information and texture information of the current feature region; determining a first state value of the current cycle according to the current observation information and a first state value of the previous cycle, the first state value of the previous cycle being determined by observation information extracted by the previous cycle; determining the first weight value according to the first state value of the current cycle; The effective channel attention mechanism is used to assign a second weight value to the feature region, including: performing global average pooling on the feature region to obtain a feature vector representing the feature region; performing a fast one-dimensional convolution calculation on the feature vector to capture local cross-channel interaction information, a convolution size of the fast one-dimensional convolution calculation being proportional to a channel dimension of the global average pooling, a coverage range of the local cross-channel interaction information being determined by the convolution size of the fast one-dimensional convolution calculation; determining channel attention according to the local cross-channel interaction information; assigning the second weight value to the feature region based on the channel attention.

7. A computer-readable storage medium storing a computer program, wherein the computer program comprises the following steps of: The computer program is executed by a processor to implement the following method: Collecting multiple images of multiple body parts of a child, the multiple images at least including an oral cavity image of the child; If the opening and closing degree of the child's mouth in the oral cavity image is greater than a preset threshold, a detection model improved based on a YOLOv5 model is used to perform disease detection on the multiple images, and a detection result is output; The detection model includes an input layer, a bone layer, a neck layer, and an output layer, the neck layer includes a feature fusion module, and the feature fusion module is used to divide the multiple images into multiple feature regions respectively and assign weight values to each feature region, and the weight values of each feature region are not completely the same; The feature fusion module includes a cyclic focusing mechanism and an effective channel attention mechanism, and the feature fusion module is used to assign weight values to each feature region, specifically including: For any feature region, the cyclic focusing mechanism is used to assign a first weight value to the feature region; The effective channel attention mechanism is used to assign a second weight value to the feature region; The weight value of the feature region is determined according to the first weight value and the second weight value; The cyclic focusing mechanism is used to assign a first weight value to the feature region, including: determining a current observation position of a current feature region; Based on the current observation position, the current observation information of the current feature region is cyclically extracted, and the current observation information includes position information and texture information of the current feature region; A first state value of the current cycle is determined according to the current observation information and a first state value of a previous cycle, and the first state value of the previous cycle is determined by observation information extracted by the previous cycle; The first weight value is determined according to the first state value of the current cycle; The effective channel attention mechanism is used to assign a second weight value to the feature region, including: performing global average pooling on the feature region to obtain a feature vector for representing the feature region; performing a fast one-dimensional convolution calculation on the feature vector to capture local cross-channel interaction information, a convolution size of the fast one-dimensional convolution calculation being proportional to a channel dimension of the global average pooling, a coverage range of the local cross-channel interaction information being determined by the convolution size of the fast one-dimensional convolution calculation; determining channel attention according to the local cross-channel interaction information; assigning the second weight value to the feature region based on the channel attention.

Citation Information

Patent Citations

  • Lithium battery defect detection method based on improved YOLOv4

    CN114049313A

  • Lymph node CT detection system employing recurrent spatio-temporal attention mechanism

    WO2020258611A1