Focus area identification method, system and device based on mobile terminal and medium
The abbreviation image is identified through multiple repeated network models on the mobile terminal, which solves the problems of cumbersome diagnosis process and communication interruption affecting efficiency in the prior art, and achieves efficient lesion recognition and diagnosis.
Patent Information
- Application Number
- CN202510450903.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-13
Smart Images

Figure CN119992074A_ABST
Abstract
Description
Background Art
[0002] The anterior segment of the eye is a general term for structures such as the cornea, anterior chamber and lens. The normal morphology and function of the anterior segment of the eye are the prerequisites for eye health. Anterior segment diseases such as cataracts, pterygium and corneal lesions will seriously affect the user's visual function and even cause blindness.
[0003] In order to facilitate diagnosis of users, the currently commonly used method is for medical personnel to use a digital slit lamp microscope to obtain an image of the user's anterior segment of the eye, transmit the image of the user's anterior segment of the eye to the cloud, call a preset image recognition model in the cloud for identification, and then transmit the identified image to the medical personnel's equipment for medical personnel to view and perform screening diagnosis.
[0004] However, the commonly used methods currently have the following technical problems: each operation requires medical staff to operate the imager to collect images and transmit them to the cloud for identification, and then the medical staff can perform screening and diagnosis based on the identification results in the cloud; the whole process is not only cumbersome and inefficient, but once the communication is interrupted, data cannot be sent and received with the cloud, which further affects the operation of medical staff and reduces processing efficiency. Summary of the invention
[0005] The present invention proposes a lesion area recognition method, system, device and medium based on a mobile terminal, which can solve the problems of cumbersome operation in the prior art and the possibility of failure to recognize due to communication interruption.
[0006] A first aspect of an embodiment of the present invention provides a lesion area recognition method based on a mobile terminal, the method being applicable to a mobile terminal, and the method comprising: Acquire an anterior segment image of the user, wherein the anterior segment image is an image extracted from a real-time video of the anterior segment of the user's eye after the real-time video of the anterior segment of the user's eye is captured by a mobile terminal; The built-in target recognition model is called to perform recognition processing on the anterior segment of the eye image to obtain the lesion area so that medical personnel can view the lesion area. The target recognition model is a model constructed by multiple repeated networks.
[0007] In combination with the first aspect, in one implementation, the target recognition model includes: an input layer, a backbone network layer, a neck layer, and a detection head layer connected in sequence; The backbone network layer is composed of a plurality of multi-branch repeated network stacks.
[0008] In combination with the first aspect, in one implementation, calling a built-in target recognition model to perform recognition processing on the anterior segment image to obtain a lesion area includes: Preprocessing the anterior segment image using the input layer to obtain a processed image, wherein the preprocessing includes: normalization processing, cropping processing, horizontal flipping processing, scaling processing and color dithering processing; Extracting images corresponding to features at different semantic levels from the processed image using the backbone network layer to obtain feature images; Calling the neck layer to fuse multiple feature images to obtain a fused image; The detection head layer is used to perform identification according to the feature points of the fused image to obtain the lesion area.
[0009] In combination with the first aspect, in one implementation, the step of concatenating the convolution features in the channel dimension to form a fusion feature includes: The fusion feature is calculated using the following formula: ; Among them, Concat(X) represents the concatenation operation in the channel dimension.
[0010] In combination with the first aspect, in one implementation, performing a residual connection on the second convolution feature and the initial feature to obtain a connected image includes: The following residual connection is used: ; Where Y is the concatenated image output by multiple repetitive networks.
[0011] In combination with the first aspect, in one implementation, extracting features of different semantic levels from the processed image using the backbone network layer to obtain a feature image includes: Using the multi-branch repeated network of the backbone network layer to perform feature segmentation on the processed image to obtain initial features and processed features; Using the convolutional layers of the multiple repeated networks of the backbone network layer to perform convolution operations and grouping operations on the processed features to obtain first convolutional features; After concatenating the convolution features in the channel dimension to form a fusion feature, a convolution operation is performed on the fusion feature to obtain a second convolution feature: After performing residual connection on the second convolution feature and the initial feature to obtain a connected image, the connected image is resized to obtain a feature image.
[0012] In combination with the first aspect, in one implementation, the step of calling the neck layer to fuse the plurality of feature images to obtain a fused image includes: The FPN network of the neck layer is called to fuse multiple feature images, and the PANet network of the neck layer is called to fuse multiple feature images to obtain a fused image with multi-scale features and bidirectional fusion.
[0013] In combination with the first aspect, in one implementation, the using the detection head layer to identify according to the feature points of the fused image to obtain the lesion area includes: After adjusting the number of channels of the fused image to a preset number of channels by using the convolution layer of the detection head layer, the category probability value is calculated according to the feature points of the adjusted fused image; A lesion area in the fused image is determined based on the category probability value.
[0014] In combination with the first aspect, in one implementation, obtaining an image of anterior segment of the eye of the user includes: Acquire real-time video, where the real-time video is obtained by zooming the lens of the mobile terminal to a focal length that can clearly observe the entire single eye of the user and capturing the video using an optical sensor; After the real-time video is resized to obtain a target video, an image corresponding to the anterior segment of the eye of the user is extracted from the target video to obtain an anterior segment image.
[0015] A second aspect of an embodiment of the present invention provides a lesion area recognition system based on a mobile terminal, the system comprising: Anterior segment image acquisition module, used to acquire real-time video of the user's anterior segment; Image processing module, used to resize real-time video; The target recognition module is used to identify the lesion area from the real-time video; A data processing module, used for performing Boolean logic processing on the recognition result of the target recognition module; The display feedback module is used to display the image output by the image processing module and the recognition result output by the data processing module.
[0016] Compared with the prior art, the embodiments of the present invention provide a method, system, device and medium for identifying a lesion area based on a mobile terminal, and its beneficial effect is that: the present invention can obtain the user's anterior segment image through the mobile terminal, and then call the built-in target recognition model to identify and process the anterior segment image to obtain the lesion area, so that medical personnel can view the lesion area, and the target recognition model is a model constructed by multiple repeated networks. The present invention can collect the user's anterior segment image at any time and anywhere through the mobile terminal, and then identify and process the anterior segment image through the built-in target recognition model to obtain the lesion area. The mobile terminal calls the built-in model for image processing, which can simplify the identification process and improve the processing efficiency. Moreover, the entire processing method allows the mobile terminal to operate offline, which can avoid the situation where it cannot be identified due to communication interruption, and does not need to rely on doctors to process, further improving the processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a flowchart of a method for identifying a lesion area based on a mobile terminal provided by an embodiment of the present invention; Figure 2 This is an operation flow chart of a lesion area identification method based on a mobile terminal provided by an embodiment of the present invention; Figure 3 It is a structural schematic diagram of a lesion area identification device based on a mobile terminal provided by an embodiment of the present invention; Figure 4 It is a structural schematic diagram of a lesion area recognition system based on a mobile terminal provided in one embodiment of the present invention. DETAILED DESCRIPTION
[0018] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0019] The anterior segment of the eye is a general term for structures such as the cornea, anterior chamber and lens. The normal morphology and function of the anterior segment of the eye are the prerequisites for eye health. Anterior segment diseases such as cataracts, pterygium and corneal lesions will seriously affect the user's visual function and even cause blindness.
[0020] In order to diagnose the user, the currently commonly used method is for medical personnel to use a digital slit lamp microscope to obtain images of the user's anterior segment and fundus, transmit the images of the user's anterior segment and fundus to the cloud, call the preset image recognition model in the cloud for recognition, and then transmit the recognized images to the medical personnel's equipment for medical personnel to view and perform screening diagnosis.
[0021] However, the commonly used methods currently have the following technical problems: each operation requires medical staff to operate the imager to collect images and transmit them to the cloud for identification, and then the medical staff can perform screening and diagnosis based on the identification results in the cloud; the whole process is not only cumbersome and inefficient, but once the communication is interrupted, data cannot be sent and received with the cloud, which further affects the operation of medical staff and reduces processing efficiency.
[0022] In order to solve the above problems, a lesion area identification method, system, device and medium based on a mobile terminal provided in an embodiment of the present application will be introduced and explained in detail through the following specific embodiments.
[0023] Reference Figure 1 , showing a flow chart of a lesion area identification method based on a mobile terminal provided by an embodiment of the present invention.
[0024] The method is applicable to a mobile terminal. The mobile terminal involved in the present invention has a central processing unit, a memory, an input component (keyboard, camera), an output component (display screen), and a diffuser in terms of hardware. Its operating system can be Windows, Android, HarmonyOS, IOS, etc.
[0025] The mobile terminal is not limited to smart phones and tablet computers. Users can directly use the mobile terminal to collect images anywhere without having to collect images on the instrument, which can simplify the operation and allow users to operate by themselves. In addition, the mobile terminal can obtain eye information through a lens and an optical sensor in conjunction with a flash.
[0026] In addition, since mobile terminals can perform real-time recognition, they can also provide real-time feedback on recognition results for medical staff to view, making it easier for medical staff to conduct intelligent diagnosis.
[0027] Wherein, as an example, the lesion area recognition method based on the mobile terminal may include: S11. Acquire an anterior segment image of the user, where the anterior segment image is an image extracted from a real-time video of the anterior segment of the user's eye after the real-time video is taken by a mobile terminal.
[0028] In one embodiment, an image of the user's anterior segment of the eye can be obtained. The image of the anterior segment of the eye can be an image extracted from a real-time video taken by medical personnel using a mobile terminal; or it can be an image extracted from a real-time video taken by the user at home using a mobile terminal.
[0029] In one operation mode, the doctor can directly use the mobile terminal to shoot video. The mobile terminal can shoot a real-time video of the user's anterior segment of the eye, and then extract a number of images containing the user's anterior segment of the eye from the real-time video of the user's anterior segment of the eye to obtain an anterior segment image.
[0030] Since the anterior segment image is obtained by taking pictures with a mobile terminal, and the mobile terminal may be unstable when taking pictures, the sizes of the anterior segment images extracted subsequently are different. In order to use images of uniform size for recognition processing to improve the recognition accuracy, as an example, the acquisition of the anterior segment image of the user may include the following sub-steps: S111. Acquire real-time video, where the real-time video is obtained by zooming the lens of the mobile terminal to a focal length that can clearly observe the entire single eye of the user and capturing the video using an optical sensor.
[0031] S112, scaling the real-time video to obtain a target video, extracting an image corresponding to the anterior segment of the eye of the user from the target video to obtain an anterior segment image.
[0032] Specifically, the mobile terminal may be composed of a lens, an optical sensor, and a flash, and the mobile terminal is used to obtain real-time video of the user's anterior segment. During actual acquisition, the flash of the mobile terminal may be turned on to illuminate the user's face, and the lens of the mobile terminal may be zoomed to a focal length that can clearly observe the user's entire single eye, and then the optical sensor may be used to acquire video through the lens to obtain real-time video.
[0033] After acquiring the real-time video, the mobile terminal can scale the real-time video, unify the input size of the image, and then extract the image containing the anterior segment of the eye from the video to obtain the anterior segment of the eye image.
[0034] S12, calling a built-in target recognition model to perform recognition processing on the anterior segment image to obtain a lesion area for medical personnel to check the lesion area, wherein the target recognition model is a model constructed by multiple repeated networks.
[0035] After acquiring the anterior segment image, the target recognition model built into the mobile terminal can be called to identify and process the anterior segment image to obtain the lesion area. After the lesion area is identified, the mobile terminal can display the lesion area for medical personnel to view, so that the medical personnel can diagnose the user based on the identified lesion area. The target recognition model is a model constructed by multiple repeated networks.
[0036] In one embodiment, a target recognition model can be pre-set in a mobile terminal for recognition. Since the model is built into the mobile terminal, the mobile terminal can be used offline, which can avoid the situation where recognition cannot be performed due to interruption of communication with the cloud, thereby improving operational efficiency and recognition efficiency.
[0037] In one embodiment, the target recognition model includes: an input layer (Input), a backbone network layer (Backbone), a neck layer (Neck) and a detection head layer (Head) connected in sequence; The backbone network layer is composed of a plurality of multi-branch repeated network stacks.
[0038] The backbone network layer can be the core part of the target recognition model, responsible for extracting multi-level features from the input image. The backbone network layer can introduce multiple repeat networks (MRN, Multiple repeat network), an efficient hierarchical aggregation method, to achieve feature reuse and fusion.
[0039] This backbone network layer can divide the input feature map into multiple parallel convolution branches, each of which may contain a different number of convolution layers to capture features of different scales and complexities. The convolution operation of each branch usually uses a 3×3 convolution kernel with a stride of 1 and a padding of 1 to keep the size of the feature map unchanged.
[0040] The number of output channels of each branch is set to 64 according to design requirements. After completing the parallel convolution, the output feature maps of each branch are spliced in the channel dimension to form a feature map with a higher number of channels. For example, if there are four branches and the number of output channels of each branch is 64, the number of channels of the spliced feature map is 256. In order to control the number of parameters and the amount of calculation of the model, the spliced feature map is usually passed through a 1×1 convolution layer with a step size of 1 and a padding of 0 to adjust the number of channels to the required setting of 128. The role of the 1×1 convolution layer is to fuse channel information while reducing the number of parameters.
[0041] The backbone network layer realizes layer-by-layer abstraction of feature maps by stacking and downsampling multiple times of multiple repeated networks. Downsampling is usually achieved through a 3×3 convolutional layer with a stride of 2. The feature map size is halved after each downsampling, while the number of channels is appropriately increased. Therefore, the feature map size changes from 512×512 to 256×256, 128×128, 64×64, 32×32, and 16×16, and the number of channels increases from 2 to 64, 128, 256, 512, and 1024. In this way, the backbone network layer can extract features of different scales and semantic levels, providing rich information for subsequent target detection.
[0042] As an example, calling a built-in target recognition model to perform recognition processing on the anterior segment image to obtain the lesion area may include the following sub-steps: S121, using the input layer to preprocess the anterior segment image to obtain a processed image, wherein the preprocessing includes: normalization processing, data enhancement processing, cropping processing, horizontal flipping processing, scaling processing and color dithering processing.
[0043] In one operation mode, the input layer can receive a color anterior segment image of the eye with a size of 512×512×3, where 512×512 represents the width and height of the image, and 3 represents the RGB three channels. In order to ensure the stability and accuracy of the model, the input image needs to be preprocessed. The preprocessing may include: normalization, cropping, horizontal flipping, scaling, and color dithering.
[0044] Specifically, normalization can normalize the pixel values of the anterior segment image to the interval [0, 1] to eliminate the influence of the numerical scale on model training. Data enhancement techniques such as random cropping, horizontal flipping, random scaling, and color jittering can be applied to increase the diversity of training data and prevent model overfitting. The preprocessed image is used as input and directly sent to the backbone network layer for feature extraction.
[0045] S122. Using the backbone network layer, extract images corresponding to features at different semantic levels from the processed image to obtain a feature image.
[0046] In one embodiment, the backbone network layer may be directly called to extract images corresponding to features at different semantic levels from the processed image, and a feature image may be obtained.
[0047] Since the backbone network layer can divide the input feature map into multiple parallel convolution branches, each branch may contain a different number of convolution layers, features of different scales and complexities can be obtained.
[0048] Wherein, as an example, the step of extracting features of different semantic levels from the processed image using the backbone network layer to obtain a feature image may include the following sub-steps: S1221. Use the multi-branch repeated network of the backbone network layer to perform feature division on the processed image to obtain initial features and processed features.
[0049] In the specific operation, it can be assumed that the input processing image is X0, and its size is H×W×C in , where H and W represent the height and width of the feature map respectively, and C in is the number of channels.
[0050] The multiple repeated networks of the backbone network layer can first perform feature division on the input processed image, and then perform feature replication after division to achieve feature reuse.
[0051] The specific steps are as follows: Feature division can divide the input processing image X0 into two parts to obtain initial features and processing features. Among them, a part of the initial features can be directly used as the input of subsequent feature fusion, and the other part of the processing features can be used for further feature extraction.
[0052] In one embodiment, the initial feature X direct And process feature X processed It can be shown as follows: ; ; Processing feature X processed Can be reused to process feature X processedIt can be used multiple times to extract deeper features through multi-layer convolution operations.
[0053] S1222. Use the convolutional layers of the multiple repeated networks of the backbone network layer to perform convolution operations and grouping operations on the processed features to obtain first convolutional features.
[0054] The convolutional layers of the multiple branches of the backbone network layer are used to process the feature X processed Perform convolution operation to process feature X processed After a series of parallel convolutional layers complete the convolution operation, each branch may contain a different number of convolutional layers to capture diverse features. Suppose there are n branches, and the specific operation of each branch is as follows: The convolution operation of the i-th branch: ; Among them, Conv (i) (X) represents the convolution operation of the i-th branch, which may contain one or more convolutional layers.
[0055] The convolutional layers of the multiple branches of the backbone network layer are used to process the feature X processed A grouping operation is performed, which may be an application of grouped convolution.
[0056] In order to reduce computational complexity, multi-branch repeated networks use grouped convolution in convolution operations. Assuming that the number of groups in each convolution layer is G, the convolution operation is only performed within each group. Specifically, the grouped convolution formula can be shown as follows: For the g-th group of input feature maps (g=1,2,...,G): ; Among them, Xprocessed,g is the processed feature X of the input processed In the g-th group of features, Conv g(i) (X,g) is the corresponding convolution operation, F i,g is the first convolution feature.
[0057] S1223. After concatenating the convolution features in the channel dimension to form a fusion feature, a convolution operation is performed on the fusion feature to obtain a second convolution feature.
[0058] The first convolution features output by all branches can be concatenated in the channel dimension to form a fused second convolution feature, and the second convolution feature can be a feature map.
[0059] In one embodiment, the feature concatenation formula may be as follows: ; Among them, Concat(X) represents the concatenation operation in the channel dimension.
[0060] S1224. After performing residual connection on the second convolutional feature and the initial feature to obtain a connected image, resize the connected image to obtain a feature image.
[0061] The second convolution feature Fconcat after fusion has a high number of channels. In order to integrate information and control the number of channels, a 1×1 convolution layer is applied, where the 1×1 convolution formula can be described as follows: ; Among them, Conv 1×1 (X) represents a 1×1 convolution operation.
[0062] In one embodiment, the second convolution feature and the initial feature can be residually connected to obtain a connected image, wherein the residual connection can directly transfer the initial feature X direct A residual connection is performed with the integrated second convolution feature Fconcat to achieve feature fusion and information retention.
[0063] Among them, the residual connection formula can be described as follows: Where Y is the concatenated image output by multiple repetitive networks.
[0064] In order to make the size of the output image the same as the size of the input image, it is H×W×C out , where C out is the number of output channels, which is consistent with the number of output channels of the 1×1 convolutional layer. The connected image can be resized to obtain the feature image.
[0065] S123, calling the neck layer to fuse the multiple feature images to obtain a fused image.
[0066] In one embodiment, since there are multiple features at different semantic levels of the processed image, there are also multiple corresponding feature images. The neck layer can be used to fuse multiple feature images to obtain a fused image.
[0067] In one embodiment, the step of calling the neck layer to fuse the plurality of feature images to obtain a fused image may include the following sub-steps: S1231, calling the FPN network of the neck layer to fuse the multiple feature images, and calling the PANet network of the neck layer to fuse the multiple feature images, to obtain a fused image with multi-scale features and bidirectional fusion.
[0068] Specifically, the neck layer combines the FPN network (Feature Pyramid Network) and the PANet network (Path Aggregation Network), has the advantages of the FPN network and the PANet network, and forms an efficient feature fusion module.
[0069] The fusion process of the FPN network starts with the feature image of the high layer of the backbone network layer, that is, the feature image of size 16×16×1024. The size of the feature image is enlarged to 32×32×1024 through upsampling operations (for example, using the nearest neighbor interpolation method with an upsampling factor of 2). Then, the upsampled feature image is fused with the feature image of the corresponding scale (32×32×512) in the backbone network layer. The fusion method is usually to splice in the channel dimension to obtain a feature image of size 32×32×1536. In order to control the number of channels and the amount of calculation, the fused feature image passes through a 1×1 convolution layer, and the number of channels is adjusted to 512. Next, the fused feature image is upsampled again to obtain a feature image of size 64×64×512, and the same fusion and channel adjustment are performed with the feature image of the low layer of the backbone network layer (64×64×256), and finally a feature image of size 64×64×256 is obtained.
[0070] The fusion process of the PANet network starts with the low-level feature images and performs bottom-up feature fusion. First, the 64×64×256 feature image is processed, downsampled to 32×32×256 through a 3×3 convolution layer with a step size of 2, and fused with the corresponding feature image of the FPN network. The fused feature image is then downsampled to 16×16×512 through a similar operation and fused with the feature image of a higher layer.
[0071] By combining the FPN network with the PANet network, the neck layer structure realizes the bidirectional fusion of multi-scale features, which not only utilizes the semantic information of high-level features, but also retains the detailed information of low-level features. This structure enables the model to have good performance when detecting objects of different sizes.
[0072] S124: Utilize the detection head layer to identify the feature points of the fused image to obtain a lesion area.
[0073] In one embodiment, the fused image may be input into a detection head layer, and the detection head layer may identify the lesion location in the image based on feature points of the fused image, thereby obtaining the lesion area in the image.
[0074] Wherein, as an example, the step of using the detection head layer to identify the feature points of the fused image to obtain the lesion area may include the following sub-steps: S1241. After adjusting the number of channels of the fused image to a preset number of channels using the convolutional layer of the detection head layer, calculate the category probability value according to the feature points of the adjusted fused image.
[0075] S1242: Determine a lesion region in the fused image based on the category probability value.
[0076] Specifically, the detection head layer can determine the classification and location prediction of the target lesion based on the fused image.
[0077] On the fused image of each scale, it first passes through a 3×3 convolution layer with a stride of 1 and padding of 1 to further extract local features. Then, a 1×1 convolution layer with a stride of 1 and padding of 0 is used to adjust the number of channels to the number of channels required for prediction, that is, 30.
[0078] This number comes from the number of prediction parameters for each feature point, calculated as B×(5+C), where B is the number of anchor boxes for each feature point in the fused image, set to 3, and C is the number of categories, which can be 5 here. During the prediction process, each feature point outputs a vector containing 10 elements for each anchor box, including 4 bounding box regression parameters (center coordinates x, y, width w, height h), 1 target confidence, and 5 category probabilities. The probability value of the lesion category is then screened, and whether this area is a lesion area is determined based on the size of the probability value.
[0079] For example, if the probability value is greater than a preset value, the region is determined to be a lesion region, and the lesion region in the image can be subsequently displayed to medical personnel.
[0080] Optionally, the prediction results of feature maps of different scales can be integrated and post-processing methods such as non-maximum suppression (NMS) can be used to remove duplicate and low-confidence detection boxes, and finally output the category and precise location of the target.
[0081] In addition, after the lesion area is identified, three Boolean logic processes can be performed on the identified lesion area to determine whether the human eye category is identified, whether the lesion category is identified, and whether the lesion rectangle is included in the human eye rectangle. If the above three logic processes are all true values, the identified lesion area will be displayed on the mobile terminal for the user to view.
[0082] Reference Figure 2 , shows an operation flow chart of a lesion area identification method based on a mobile terminal provided by an embodiment of the present invention.
[0083] Specifically, the operation of the lesion area recognition method based on the mobile terminal may include the following steps: The first step is to obtain the user's identification information, including the user's anterior eye image.
[0084] The second step is to determine whether the user's identification information includes human eyes.
[0085] In the third step, if the user's identification information includes human eyes, the lesion information is identified, including the lesion area.
[0086] The fourth step is to show the lesion area to the user.
[0087] In this embodiment, the embodiment of the present invention provides a method for identifying a lesion area based on a mobile terminal, and its beneficial effect is that the present invention can obtain an image of the user's anterior segment of the eye through a mobile terminal, and then call a built-in target recognition model to identify and process the image of the anterior segment of the eye to obtain a lesion area, so that medical personnel can view the lesion area, and the target recognition model is a model constructed by multiple repeated networks. The present invention allows medical personnel to collect an image of the user's anterior segment of the eye through a mobile terminal anytime and anywhere, and then identify and process the image of the anterior segment of the eye through a built-in target recognition model to obtain a lesion area, which can simplify the processing flow of identification and improve processing efficiency. Moreover, the entire processing method allows the mobile terminal to operate offline, which can avoid the situation where identification cannot be performed due to communication interruption, and does not need to rely on doctors to process, further improving processing efficiency.
[0088] The embodiment of the present invention also provides a lesion area recognition device based on a mobile terminal, see Figure 3 , showing a schematic structural diagram of a lesion area identification device based on a mobile terminal provided by an embodiment of the present invention.
[0089] The device is applicable to a mobile terminal, wherein, as an example, the lesion area identification device based on the mobile terminal may include: An acquisition module 310 is used to acquire an anterior segment image of the user, where the anterior segment image is an image extracted from a real-time video of the anterior segment of the user's eye after the real-time video of the anterior segment of the user's eye is captured by a mobile terminal; The lesion recognition module 320 is used to call the built-in target recognition model to perform recognition processing on the anterior segment image to obtain the lesion area for medical personnel to view the lesion area. The target recognition model is a model constructed by multiple repeated networks.
[0090] Optionally, the target recognition model comprises: an input layer, a backbone network layer, a neck layer and a detection head layer connected in sequence; The backbone network layer is composed of a plurality of multi-branch repeated network stacks.
[0091] Optionally, calling a built-in target recognition model to perform recognition processing on the anterior segment image to obtain a lesion area includes: Preprocessing the anterior segment image using the input layer to obtain a processed image, wherein the preprocessing includes: normalization processing, cropping processing, horizontal flipping processing, scaling processing and color dithering processing; Extracting images corresponding to features at different semantic levels from the processed image using the backbone network layer to obtain feature images; Calling the neck layer to fuse multiple feature images to obtain a fused image; The detection head layer is used to perform identification according to the feature points of the fused image to obtain the lesion area.
[0092] Optionally, the extracting features of different semantic levels from the processed image using the backbone network layer to obtain a feature image includes: Using the multi-branch repeated network of the backbone network layer to perform feature segmentation on the processed image to obtain initial features and processed features; Using the convolutional layers of the multiple repeated networks of the backbone network layer to perform convolution operations and grouping operations on the processed features to obtain first convolutional features; After concatenating the convolution features in the channel dimension to form a fusion feature, a convolution operation is performed on the fusion feature to obtain a second convolution feature: After performing residual connection on the second convolutional feature and the initial feature to obtain a connected image, the connected image is resized to obtain a feature image.
[0093] Optionally, the calling of the neck layer to fuse the plurality of feature images to obtain a fused image includes: The FPN network of the neck layer is called to fuse multiple feature images, and the PANet network of the neck layer is called to fuse multiple feature images to obtain a fused image with multi-scale features and bidirectional fusion.
[0094] Optionally, the using the detection head layer to identify according to the feature points of the fused image to obtain the lesion area includes: After adjusting the number of channels of the fused image to a preset number of channels by using the convolution layer of the detection head layer, the category probability value is calculated according to the feature points of the adjusted fused image; A lesion area in the fused image is determined based on the category probability value.
[0095] Optionally, obtaining an image of the user's anterior segment of the eye includes: Acquire real-time video, where the real-time video is obtained by zooming the lens of the mobile terminal to a focal length that can clearly observe the entire single eye of the user and capturing the video using an optical sensor; After the real-time video is resized to obtain a target video, an image corresponding to the anterior segment of the eye of the user is extracted from the target video to obtain an anterior segment image.
[0096] The embodiment of the present invention also provides a lesion area recognition system based on a mobile terminal, see Figure 4 , showing a structural schematic diagram of a lesion area identification system based on a mobile terminal provided in one embodiment of the present invention.
[0097] Wherein, as an example, the lesion area recognition system based on the mobile terminal may include: The anterior segment image acquisition module is used to acquire real-time video of the anterior segment of the user's eye; the image acquisition module is composed of a lens, an optical sensor, and a flash of the mobile terminal.
[0098] Image processing module, used to resize real-time video; The target recognition module is used to identify the lesion area from the real-time video; specifically, the target recognition module is a target recognition algorithm developed based on deep learning, which can perform target recognition on the incoming video in real time, and recognize a total of 5 types of targets: human eyes, cataract lesions, pterygium lesions, corneal lesion areas and subconjunctival hemorrhage areas, and transmit the recognition information of each frame (including the target recognition category and the corresponding rectangular frame coordinates) to the data processing module.
[0099] A data processing module, used for performing Boolean logic processing on the recognition result of the target recognition module; The display feedback module is used to display the image output by the image processing module and the recognition result output by the data processing module.
[0100] Those skilled in the art can clearly understand that, for the sake of convenience and brevity of description, the specific working process of the system described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0101] Furthermore, an embodiment of the present application also provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the mobile terminal-based lesion area identification method as described in the above embodiment is implemented.
[0102] Furthermore, an embodiment of the present application also provides a computer-readable storage medium, which stores a computer-executable program, and the computer-executable program is used to enable a computer to execute the lesion area identification method based on a mobile terminal as described in the above embodiment.
[0103] It should be noted that the orientation or positional relationship indicated by the terms "upper", "lower", etc. is based on the orientation or positional relationship shown in the accompanying drawings, which is only for the convenience of describing the embodiments of the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention. When an element such as a layer, region or substrate is referred to as being "on" or "above" another element, it can be directly on the other element, or there can also be an intermediate element. On the contrary, when an element is referred to as "directly on" or "above" another element, there is no intermediate element. It should also be understood that when an element is referred to as being "under" or "below" another element, it can be directly under or below the other element, or there can also be an intermediate element. On the contrary, when an element is referred to as being "directly under" or "below" another element, there is no intermediate element. Unless otherwise clearly specified and limited, the terms "installed", "connected" and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0104] Those skilled in the art will appreciate that the embodiments of the present application may also provide computer program products. Therefore, the present application may adopt the form of complete hardware embodiments, complete software embodiments, or embodiments in combination with software and hardware. Moreover, the present application may adopt the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program codes.
[0105] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), apparatuses, and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0106] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0107] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0108] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A lesion area recognition method based on a mobile terminal, characterized in that: The method is applicable to a mobile terminal, and the method comprises: Acquire an anterior segment image of the user, wherein the anterior segment image is an image extracted from a real-time video of the anterior segment of the user's eye after the real-time video of the anterior segment of the user's eye is captured by a mobile terminal; Calling a built-in target recognition model to perform recognition processing on the anterior segment image to obtain a lesion area for medical personnel to view the lesion area, wherein the target recognition model is a model constructed by multiple repeated networks; The target recognition model includes: an input layer, a backbone network layer, a neck layer and a detection head layer connected in sequence; The backbone network layer is composed of a plurality of multi-branch repeated network stacks; The calling of the built-in target recognition model to perform recognition processing on the anterior segment image to obtain the lesion area includes: Preprocessing the anterior segment image using the input layer to obtain a processed image, wherein the preprocessing includes: normalization processing, cropping processing, horizontal flipping processing, scaling processing and color dithering processing; Extracting images corresponding to features at different semantic levels from the processed image using the backbone network layer to obtain feature images; Calling the neck layer to fuse multiple feature images to obtain a fused image; The detection head layer is used to perform identification according to the feature points of the fused image to obtain the lesion area.
2. The method for identifying a lesion region based on a mobile terminal according to claim 1, characterized in that: The step of extracting features of different semantic levels from the processed image using the backbone network layer to obtain a feature image includes: Using the multi-branch repeated network of the backbone network layer to perform feature segmentation on the processed image to obtain initial features and processed features; Using the convolutional layers of the multiple repeated networks of the backbone network layer to perform convolution operations and grouping operations on the processed features to obtain first convolutional features; After concatenating the convolution features in the channel dimension to form a fusion feature, a convolution operation is performed on the fusion feature to obtain a second convolution feature: After performing residual connection on the second convolution feature and the initial feature to obtain a connected image, the connected image is resized to obtain a feature image.
3. The method for identifying a lesion region based on a mobile terminal according to claim 2, characterized in that: The step of splicing the convolution features in the channel dimension to form a fusion feature includes: The fusion feature is calculated using the following formula: ; Among them, Concat(X) represents the concatenation operation in the channel dimension.
4. The method for identifying a lesion region based on a mobile terminal according to claim 2, characterized in that: The performing residual connection on the second convolution feature and the initial feature to obtain a connected image includes: The following residual connection is used: ; Where Y is the concatenated image output by multiple repetitive networks.
5. The method for identifying lesion areas based on a mobile terminal according to claim 1, characterized in that: The calling of the neck layer to fuse the plurality of feature images to obtain a fused image includes: The FPN network of the neck layer is called to fuse multiple feature images, and the PANet network of the neck layer is called to fuse multiple feature images to obtain a fused image with multi-scale features and bidirectional fusion.
6. The method for identifying a lesion region based on a mobile terminal according to claim 1, characterized in that: The step of using the detection head layer to identify the feature points of the fused image to obtain the lesion area includes: After adjusting the number of channels of the fused image to a preset number of channels by using the convolution layer of the detection head layer, the category probability value is calculated according to the feature points of the adjusted fused image; A lesion area in the fused image is determined based on the category probability value.
7. The method for identifying a lesion region based on a mobile terminal according to any one of claims 1 to 6, characterized in that: The step of obtaining an image of the anterior segment of the eye of the user comprises: Acquire real-time video, where the real-time video is obtained by zooming the lens of the mobile terminal to a focal length that can clearly observe the entire single eye of the user and capturing the video using an optical sensor; After the real-time video is resized to obtain a target video, an image corresponding to the anterior segment of the eye of the user is extracted from the target video to obtain an anterior segment image.
8. A lesion area recognition system based on a mobile terminal, characterized in that: The system comprises: Anterior segment image acquisition module, used to acquire real-time video of the user's anterior segment; Image processing module, used to resize real-time video; The target recognition module is used to identify the lesion area from the real-time video; A data processing module, used for performing Boolean logic processing on the recognition result of the target recognition module; The display feedback module is used to display the image output by the image processing module and the recognition result output by the data processing module.
9. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the mobile terminal-based lesion area identification method as described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer-executable program, and the computer-executable program is used to enable a computer to execute the lesion area identification method based on a mobile terminal as described in any one of claims 1-7.
Citation Information
Patent Citations
Stomach focus recognition model training method and stomach focus recognition method
CN114140651A
Defect detection method and device, computer equipment and storage medium
CN116977239A
Esophageal cancer image recognition and classification method, system and equipment and medium
CN117218129A
Endoscope image part classification and focus identification method, device, equipment and medium
CN118097379A